Posted on: 11/08/2026
Job Description :
Build and scale data pipelines for NyaayOS, processing large volumes of messy legal documents. Work with Python, AWS, PostgreSQL, OCR and async workflows to build reliable ingestion, document processing, data-quality and provenance systems.
What we're looking for :
- 4+ yrs in data/backend engineering.
- Strong Python, AWS, PostgreSQL & production data pipelines.
- Hands-on OCR/document processing and async workflows using Airflow, Celery, Dagster or Temporal.
Key Responsibilities :
- Design and implement scalable ETL pipelines to ingest and transform complex, unstructured data sets into structured formats for downstream analytics.
- Develop and maintain high-performance data processing workflows using Airflow to ensure timely and reliable data availability.
- Integrate OCR technologies into existing data pipelines to automate the extraction of information from varied document formats.
- Optimize database performance and schema design within PostgreSQL to support high-concurrency read and write operations.
- Build and manage distributed task queues using Celery and Temporal to handle asynchronous processing and fault-tolerant background jobs.
- Deploy and manage cloud-native data infrastructure on AWS, ensuring cost-efficiency and high availability of data services.
- Collaborate with engineering leadership to define data architecture standards and improve overall system reliability.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1662027