HamburgerMenu
hirist

Senior Data Scientist - Machine Learning

InfoTrellis India Pvt Ltd
10 - 12 Years
Multiple Locations

Posted on: 30/06/2026

Job Description

Senior Data Scientist

We are looking for a Senior Data Scientist with deep expertise in machine learning, intelligent document processing, and cloud-native AI pipelines. You will lead the design and delivery of production-grade AI/ML components within a governed, human-in-the-loop automation platform, working across the full spectrum from raw data ingestion to model deployment and continuous improvement.

Key Responsibilities :

- Architect and implement multi-tier document extraction pipelines combining deterministic parsing, OCR, and large language model (LLM) inference, applying the right technique for each document category.

- Design and train document classification models (e.g., using Amazon SageMaker) for routing documents to appropriate extraction strategies.

- Develop and optimize prompt engineering patterns for LLM-based extraction (Amazon Bedrock), including confidence scoring, source-region anchoring, and hallucination detection strategies.

- Build and maintain data validation frameworks: schema validation, cross-field consistency checks, and configurable confidence thresholds per field and document type.

- Design canonical data models and normalization pipelines that consolidate heterogeneous source data into a unified schema with full field-level lineage.

- Lead the design of agentic AI components (extraction orchestration, layout detection, exception analysis) with clearly defined capability envelopes, escalation logic, and human approval gates.

- Define and track extraction accuracy metrics; drive continuous improvement loops using structured human feedback.

- Collaborate with solution architects to ensure ML components integrate cleanly with serverless AWS infrastructure (Lambda, Step Functions, S3, Glue, DynamoDB, EventBridge).

- Mentor junior team members and establish best practices for model governance, change management, and reproducibility.

Machine Learning & AI :

- Document classification using supervised ML (SageMaker or equivalent)

- LLM prompt engineering for structured data extraction (Bedrock, OpenAI, or similar)

- OCR post-processing : table reconstruction, merged cell handling, multi-line value normalization (Amazon Textract or equivalent)

- Confidence scoring design and threshold calibration

- Cosine similarity and document fingerprinting for template matching

- Agentic AI orchestration patterns with governed fallback and escalation design

Data Engineering & Python :

- Python : pandas, openpyxl, boto3 production-grade, not notebook-only

- ETL pipeline design for structured (Excel/CSV), semi-structured (PDF tables), and unstructured document types

- Schema design and canonical data modeling

- Data quality frameworks : completeness detection, duplicate identification, cross-field validation

- SQL and relational data modeling (PostgreSQL preferred)

Cloud - AWS :

- AWS Lambda, Step Functions, S3, SageMaker, Bedrock, Textract, Glue, DynamoDB, EventBridge, SQS, KMS

- Serverless-first architecture patterns; cost-efficient compute design for seasonal/batch workloads

Governance & MLOps :

- Model registry, versioning, and change board processes

- Audit trail design: immutable lineage from source document to output

- CI/CD integration for ML pipeline components

- RBAC and data security in multi-tenant cloud environments

Experience :

- 8+ years in data science or ML engineering roles

- At least 2 production deployments involving intelligent document processing or NLP pipelines

- Demonstrated experience designing human-in-the-loop systems, not just fully automated models

Preferred Candidate Profile :

- Power BI or Amazon QuickSight for operational dashboards

- React or familiarity with audit workbench UI requirements for human-in-the-loop review queues

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...