HamburgerMenu
hirist

AI Engineer - Generative AI/LLM

Inypeople Technology
5 - 16 Years
Bangalore

Posted on: 06/08/2026

Job Description

Job Specific Duties and Responsibilities :

- End-to-end ML ownership : Drive the complete lifecycle data curation, model building, evaluation, deployment, monitoring, and retraining for both predictive and generative AI systems.

- Production-grade MLOps : Build scalable pipelines for training, CI/CD, model registry, A/B testing, drift detection, and automated retraining. Optimize inference for latency, throughput, and cost.

- LLMs and SLMs : Fine-tune and deploy open and closed models using techniques such as LoRA/QLoRA, PEFT, instruction tuning, and preference tuning (RLHF/DPO). Apply quantization and distillation where needed.

- Agentic systems : Design and productionize agentic frameworks RAG pipelines, tool/function calling, memory, planning loops, and multi-agent orchestration with appropriate guardrails and observability.

- Quality and trust : Build evaluation frameworks (offline + online, including LLM-as-judge and red-teaming). Diagnose and mitigate hallucinations, bias, and drift.

- Rapid innovation : Track SOTA research, prototype quickly, and showcase work through demos and tech talks to internal stakeholders and leadership.

AI Engineer (Python, GenAI/LLMs + ML Fundamentals) :

Required Qualifications :

- 5+ years of hands-on experience as an AI/ML Engineer or Applied Scientist, with proven production deployments including at least one LLM-based or agentic system taken to production.

- Strong Python skills and solid software engineering fundamentals (version control, testing, design patterns, code reviews).

- Deep Learning & NLP : Strong grasp of transformer architectures, attention, tokenization, embeddings, and modern NLP techniques. Hands-on with PyTorch and the Hugging Face ecosystem (Transformers, PEFT, TRL, Accelerate).

- Agentic & RAG stack : Working knowledge of frameworks such as LangChain / LangGraph / LlamaIndex / CrewAI / AutoGen, plus vector stores (Pinecone, Weaviate, Qdrant, pgvector, or FAISS) and reranking strategies.

- Serving & optimization : Experience with inference servers such as vLLM, TGI, or Triton, and familiarity with quantization (GPTQ, AWQ, GGUF).

- MLOps & infra : Hands-on with tools like MLflow, Weights & Biases, Airflow, or Kubeflow; comfortable with Docker, Kubernetes, GPU workloads, and at least one major cloud (AWS / Azure / GCP).

- Soft skills : High bias for action, strong communication, ownership mindset, and intellectual curiosity.

- Nice to have : Open-source contributions, multimodal model experience, on-device SLM deployment, or familiarity with LLM security (OWASP LLM Top 10).

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...