Posted on: 17/07/2026
Job Description:
We are seeking a hands-on AI/ML Engineer to build and scale a production talent-intelligence platform that parses millions of resumes into structured, auditable data and matches them to job descriptions with explainable, evidence-backed ranking. The ideal candidate combines deep NLP/document-intelligence expertise with disciplined LLM engineering someone who treats token economics, determinism, and evaluation gates as first-class engineering concerns, not afterthoughts.
Location: Chennai Type: Fulltime
Key Responsibilities:
1. Hybrid Document Intelligence:
- Design deterministic-first extraction pipelines (regex, gazetteers, derivation rules) with LLM tiers reserved for residual fields targeting 75%+ token-cost reduction versus pure-LLM parsing.
- Solve real-world document chaos: multi-column PDF reading order, portal-export layout variants, stacked tables, OCR routing, hidden-text and keyword-stuffing detection.
2. LLM Engineering & Cost Optimization:
- Build constrained-generation integrations (response schemas, output caps, runaway/truncation handling, graceful degradation) against Gemini/OpenAI/Claude-class APIs.
- Operate batch inference at scale with per-call token accounting, cost projection, and provider abstraction.
- Design escalation ladders: deterministic ? residual prompt ? per-field micro-prompts, with confidence-gated routing.
3. Retrieval & Matching:
- Build hybrid retrieval (BM25 + dense embeddings, reciprocal-rank fusion) over vector stores (Qdrant), with cross-encoder / LLM re-ranking and grounded, quote-verified narratives.
- Implement deterministic scoring with per-tenant weights, must-have gates, and full score breakdowns for recruiter trust.
4. Evaluation & Quality Engineering:
- Own golden-set regression gates, full-corpus diff validation, LLM-as-judge audit loops, and error-taxonomy logging that feeds future fine-tuning.
- Enforce precision targets on PII fields (?98%) and date-derived arithmetic (?97%) values the LLM never computes.
5. Platform & Data Engineering:
- Ship FastAPI services with typed schemas (pydantic v2), multi-tenant isolation, API-key grants, and restricted-data serving invariants.
- Design MongoDB/S3-backed persistence for millions of documents with content-addressed dedup and GDPR deletion paths (atomic PII sub-documents).
6. Responsible AI:
- Maintain evidence-based skill scoring that is auditable by construction (NYC LL144 / EU AI Act context); no black-box rankings.
Required Skills & Qualifications:
- Technical Expertise: Python 3.11+, pydantic, FastAPI, spaCy, embeddings (sentence-transformers/MiniLM class), MongoDB, vector databases (Qdrant/pgvector/Pinecone).
- AI/ML Specialization: LLM prompt & schema engineering, constrained decoding, RAG/hybrid retrieval, NER, entity normalization & alias/discovery loops.
- Engineering Discipline: deterministic/reproducible pipelines, evaluation harnesses, regression gates, token/cost accounting, CI.
- Cloud Experience: S3-scale batch processing; AWS or GCP; Gemini/Bedrock/Anthropic API integration.
- Soft Skills: evidence-driven decision making, writing crisp technical docs, working directly with recruiter/product feedback.
- Experience: 5 to 8 years in ML/NLP or data-intensive backend engineering, with at least 2 years shipping LLM-integrated production systems.
Preferred Qualifications:
- Experience with resume/JD parsing, HR-tech, or search/ranking systems.
- SLM fine-tuning experience (Qwen/Gemma class) and building training sets from production error logs.
- Familiarity with AI hiring-regulation compliance (NYC LL144, EU AI Act) and PII/GDPR engineering.
- Cross-encoder re-ranking and LLM-as-judge evaluation experience.
Did you find something suspicious?