HamburgerMenu
hirist

Machine Learning Engineer - LLM/RAG

Spectral Consultants
7 - 12 Years
Noida

Posted on: 19/05/2026

Job Description

Description :


About the Role :


We are looking for an experienced Machine Learning Engineer specializing in Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), AI Modeling, and MLOps to design, develop, and deploy production-grade AI solutions.


The ideal candidate should have hands-on experience across the complete ML lifecycle - from data preparation and model fine-tuning to deployment, monitoring, and optimization of scalable AI systems on cloud platforms. This role demands strong expertise in generative AI, model optimization, distributed training, cloud-native deployment, and responsible AI practices.


You will work closely with product, engineering, design, and business teams to deliver innovative AI-powered solutions while mentoring junior engineers and driving technical strategy.


Key Responsibilities :


- Design and build enterprise-grade applications powered by Large Language Models (LLMs).


- Develop and optimize Retrieval-Augmented Generation (RAG) pipelines for domain-specific use cases.


- Build AI agents, prompt orchestration frameworks, and multi-step reasoning workflows.


- Implement prompt engineering strategies, prompt templates, and context management.


- Develop guardrails, hallucination detection, moderation pipelines, and AI safety mechanisms.


- Build evaluation frameworks for model performance, prompt effectiveness, and business KPIs.


- Manage the end-to-end ML lifecycle from data collection to production deployment.


- Fine-tune foundation models using techniques such as LoRA, QLoRA, PEFT, and instruction tuning.


- Conduct supervised fine-tuning and experimentation on domain-specific datasets.


- Support advanced training approaches including transfer learning and model adaptation.


- Work with structured, unstructured, and multimodal datasets for training pipelines.


- Deploy scalable AI/ML solutions on AWS cloud infrastructure.


- Build and manage deployments using :


i. AWS SageMaker


ii. Amazon Bedrock


iii. AWS Lambda


iv. EKS / ECS


v. API Gateway


vi. S3 / Glue / Athena


- Build inference APIs and model-serving pipelines for low-latency production environments.


- Implement autoscaling and cost optimization strategies for model serving.


Model Optimization & Performance Engineering :


- Optimize model training and inference performance using :


i. DeepSpeed


ii. Quantization techniques


iii. Mixed precision training


iv. GPU optimization


- Work with GPU clusters and distributed training environments.


- Optimize vector search and embedding pipelines using vector databases.


- Build scalable retrieval pipelines using embeddings and semantic search.


- Work with vector databases such as :


i. Pinecone


ii. Weaviate


iii. Chroma


iv. FAISS


v. Milvus


- Design chunking, indexing, reranking, and retrieval strategies.


- Improve search relevance, latency, and retrieval quality.


- Establish end-to-end MLOps pipelines including CI/CD for ML workflows.


- Implement model versioning, experiment tracking, and monitoring.


- Work with tools such as MLflow, Weights & Biases, or equivalent.


- Set up drift monitoring, alerting, retraining workflows, and performance dashboards.


- Ensure compliance with responsible AI, governance, explainability, and auditability requirements.


- Perform exploratory data analysis (EDA) and feature research for model development.


- Build and maintain ETL/data pipelines for ML training and inference.


- Work with big data frameworks such as Spark or Flink.


- Conduct domain-focused research to improve model performance and business outcomes.


- Collaborate with product managers, UX teams, engineering teams, and stakeholders.


- Translate business requirements into scalable ML/AI solutions.


- Mentor junior engineers and guide technical best practices.


- Drive architecture decisions, code reviews, and deployment standards.


Required Skills & Qualifications :


- 5-12 years of experience in Machine Learning / AI Engineering.


- Minimum 3+ years of experience building and deploying production ML systems.


- 1+ years of hands-on experience with LLMs, RAG systems, and generative AI applications.


- Strong programming skills in Python.


- Hands-on experience with :


i. PyTorch


ii. Hugging Face Transformers


iii. LangChain / LlamaIndex


iv. OpenAI / Anthropic / open-source LLMs


- Strong AWS cloud expertise :


i. SageMaker


ii. Bedrock


iii. Lambda


iv. ECS/EKS


v. S3


- Experience with containerization and orchestration :


i. Docker


ii. Kubernetes


- Strong knowledge of CI/CD pipelines and ML monitoring.


- Experience with vector databases and embedding models.


- Strong understanding of ML system design and production architecture.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...