Posted on: 06/08/2026
Job Specific Duties and Responsibilities :
- End-to-end ML ownership : Drive the complete lifecycle data curation, model building, evaluation, deployment, monitoring, and retraining for both predictive and generative AI systems.
- Production-grade MLOps : Build scalable pipelines for training, CI/CD, model registry, A/B testing, drift detection, and automated retraining. Optimize inference for latency, throughput, and cost.
- LLMs and SLMs : Fine-tune and deploy open and closed models using techniques such as LoRA/QLoRA, PEFT, instruction tuning, and preference tuning (RLHF/DPO). Apply quantization and distillation where needed.
- Agentic systems : Design and productionize agentic frameworks RAG pipelines, tool/function calling, memory, planning loops, and multi-agent orchestration with appropriate guardrails and observability.
- Quality and trust : Build evaluation frameworks (offline + online, including LLM-as-judge and red-teaming). Diagnose and mitigate hallucinations, bias, and drift.
- Rapid innovation : Track SOTA research, prototype quickly, and showcase work through demos and tech talks to internal stakeholders and leadership.
Required Qualifications :
- 5+ years of hands-on experience as an AI/ML Engineer or Applied Scientist, with proven production deployments including at least one LLM-based or agentic system taken to production.
- Strong Python skills and solid software engineering fundamentals (version control, testing, design patterns, code reviews).
- Deep Learning & NLP : Strong grasp of transformer architectures, attention, tokenization, embeddings, and modern NLP techniques. Hands-on with PyTorch and the Hugging Face ecosystem (Transformers, PEFT, TRL, Accelerate).
- Agentic & RAG stack : Working knowledge of frameworks such as LangChain / LangGraph / LlamaIndex / CrewAI / AutoGen, plus vector stores (Pinecone, Weaviate, Qdrant, pgvector, or FAISS) and reranking strategies.
- Serving & optimization : Experience with inference servers such as vLLM, TGI, or Triton, and familiarity with quantization (GPTQ, AWQ, GGUF).
- MLOps & infra : Hands-on with tools like MLflow, Weights & Biases, Airflow, or Kubeflow; comfortable with Docker, Kubernetes, GPU workloads, and at least one major cloud (AWS / Azure / GCP).
- Soft skills : High bias for action, strong communication, ownership mindset, and intellectual curiosity.
Education :
- B.Tech or M.Tech in Computer Science, Data Science Engineering, AI/ML Engineering, or a closely related quantitative discipline.
- Equivalent practical experience supported by a strong portfolio (open-source work, publications, or production deployments) will also be considered.
Soft Skills :
- Strong problem-solving and ownership mindset; comfortable operating in ambiguity.
- Clear communication of technical tradeoffs and experiment results to stakeholders.
- Collaborative approach with engineering, product, and data teams.
Did you find something suspicious?