HamburgerMenu
hirist

AI Data Engineer - Claims Intelligence & Data Pipelines

CRESCENDO GLOBAL LEADERSHIP HIRING INDIA PRIVATE L
2 - 10 Years
Multiple Locations

Posted on: 09/09/2026

Job Description

Job Description :

JD - AI Engineer Claims Intelligence & Data Pipelines

Total Experience: 2 - 10 years

About the Role:

The Claims platform extracts AI-driven insights from claim notes and documents to surface emerging trends for underwriters, adjusters, actuaries, and portfolio managers. We are looking for an AI Engineer with strong data pipeline experience to help operationalize and connect multiple DS-coded AI modules into production-grade systems. This is a technically deep role, heavier on AI engineering and data pipeline work - UI expertise is handled by a separate team.

What You Will Do :

AI Module Integration & Inference Pipelines :

- Integrate and adjust inference pipelines for NLP modules including document classification, entity extraction, de-identification (DEID), and LLM-based early trend detection.

- Connect DS-coded AI modules into end-to-end production workflows via Airflow DAGs on AWS EKS.

- Build and tune hybrid search pipelines combining GTE multilingual dense embeddings with GIN lexical search on Aurora PostgreSQL.

- Integrate with OpenAI-based API platform for multilingual query expansion and LLM-driven trend detection.

Document Processing & Parsing :

- Design and maintain document preprocessing pipelines that parse deeply nested JSON structures (emails with attachments, embedded PDFs) from S3/DataLake.

- Handle multilingual unstructured text (English, Spanish, Portuguese, German, Dutch, French, Italian) across 300 GB of claim notes and documents.

- Build chunking strategies and metadata extraction for downstream embedding and retrieval workflows.

Data Pipeline Engineering :

- Author and maintain Airflow DAGs for batch processing (monthly entity refresh, trend detection, DEID pipeline).

- Manage data flow across AWS services: S3, Athena, Glue, Fargate, SQS, Step Functions.

- Scale pipelines to handle 500K+ claims and hundreds of millions of text chunks.

Production Deployment & Quality :

- Deploy and version models using MLflow and Databricks.

- Manage schema evolution and migrations using Liquibase on Aurora PostgreSQL.

- Instrument pipelines with logging, monitoring, and evaluation scoring for retrieval quality.

Required Skills :

- Languages: Python (primary), SQL.

- AI / NLP: LLM API integration, multilingual embeddings (e.g., GTE), hybrid search, text classification, entity extraction, NER, PII masking.

- Data Pipelines: Apache Airflow, batch orchestration, large-scale unstructured data processing.

- Cloud & Infrastructure: AWS (S3, Athena, Glue, Fargate, EKS, SQS, Step Functions).

- Databases: PostgreSQL / Aurora, pgvector, GIN indexes, full-text search.

- ML Platform: MLflow, Databricks / Azure Databricks.

- DevOps: GitHub, CI/CD pipelines.

Nice to Have :

- Experience with RAG (Retrieval-Augmented Generation) pipeline design and evaluation.

- Familiarity with insurance claims or financial services data.

- Experience with multilingual NLP at scale.

- Knowledge of Liquibase for database schema management.

- Exposure to Runway Model Catalog or similar model deployment platforms.

Team Context :

You will work closely with Data Scientists to bridge the gap between experimental DS code and production-grade AI systems. The team follows an agile process (Jira-tracked), maintains documentation in Confluence, and operates on a modern AWS + Databricks stack.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...