Posted on: 24/09/2026
We are seeking a hands-on AI/ML Engineer to build and scale a production talent-intelligence platform that parses millions of resumes into structured, auditable data and matches them to job descriptions with explainable, evidence-backed ranking. The ideal candidate combines deep NLP/document-intelligence expertise with disciplined LLM engineering - someone who treats token economics, determinism, and evaluation gates as first-class engineering concerns, not afterthoughts.
Location : Chennai
Type : Fulltime
Key Responsibilities :
Hybrid Document Intelligence :
- Design deterministic-first extraction pipelines (regex, gazetteers, derivation rules) with LLM tiers reserved for residual fields - targeting 75%+ token-cost reduction versus pure-LLM parsing.
- Solve real-world document chaos: multi-column PDF reading order, portal-export layout variants, stacked tables, OCR routing, hidden-text and keyword-stuffing detection.
LLM Engineering & Cost Optimization :
- Build constrained-generation integrations (response schemas, output caps, runaway/truncation handling, graceful degradation) against Gemini/OpenAI/Claude-class APIs.
- Operate batch inference at scale with per-call token accounting, cost projection, and provider abstraction.
- Design escalation ladders: deterministic - residual prompt - per-field micro-prompts, with confidence-gated routing.
Retrieval & Matching :
- Build hybrid retrieval (BM25 + dense embeddings, reciprocal-rank fusion) over vector stores (Qdrant), with cross-encoder / LLM re-ranking and grounded, quote-verified narratives.
- Implement deterministic scoring with per-tenant weights, must-have gates, and full score breakdowns for recruiter trust.
Evaluation & Quality Engineering :
- Own golden-set regression gates, full-corpus diff validation, LLM-as-judge audit loops, and error-taxonomy logging that feeds future fine-tuning.
- Enforce precision targets on PII fields (>=98%) and date-derived arithmetic (>=97%) - values the LLM never computes.
Platform & Data Engineering :
- Ship FastAPI services with typed schemas (pydantic v2), multi-tenant isolation, API-key grants, and restricted-data serving invariants.
- Design MongoDB/S3-backed persistence for millions of documents with content-addressed dedup and GDPR deletion paths (atomic PII sub-documents).
Responsible AI :
- Maintain evidence-based skill scoring that is auditable by construction (NYC LL144 / EU AI Act context); no black-box rankings.
Required Skills & Qualifications :
- Technical Expertise : Python 3.11+, pydantic, FastAPI, spaCy, embeddings (sentence-transformers/MiniLM class), MongoDB, vector databases (Qdrant/pgvector/Pinecone).
- AI/ML Specialization : LLM prompt & schema engineering, constrained decoding, RAG/hybrid retrieval, NER, entity normalization & alias/discovery loops.
- Engineering Discipline : deterministic/reproducible pipelines, evaluation harnesses, regression gates, token/cost accounting, CI.
- Cloud Experience : S3-scale batch processing; AWS or GCP; Gemini/Bedrock/Anthropic API integration.
- Soft Skills : evidence-driven decision making, writing crisp technical docs, working directly with recruiter/product feedback.
- Experience : 5 - 8 years in ML/NLP or data-intensive backend engineering, with at least 2 years shipping LLM-integrated production systems.
Preferred Qualifications :
- Experience with resume/JD parsing, HR-tech, or search/ranking systems.
- SLM fine-tuning experience (Qwen/Gemma class) and building training sets from production error logs.
- Familiarity with AI hiring-regulation compliance (NYC LL144, EU AI Act) and PII/GDPR engineering.
- Cross-encoder re-ranking and LLM-as-judge evaluation experience.
Did you find something suspicious?