Posted on: 19/05/2026
Description :
About the Role :
We are looking for an experienced Machine Learning Engineer specializing in Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), AI Modeling, and MLOps to design, develop, and deploy production-grade AI solutions.
The ideal candidate should have hands-on experience across the complete ML lifecycle - from data preparation and model fine-tuning to deployment, monitoring, and optimization of scalable AI systems on cloud platforms. This role demands strong expertise in generative AI, model optimization, distributed training, cloud-native deployment, and responsible AI practices.
You will work closely with product, engineering, design, and business teams to deliver innovative AI-powered solutions while mentoring junior engineers and driving technical strategy.
Key Responsibilities :
- Design and build enterprise-grade applications powered by Large Language Models (LLMs).
- Develop and optimize Retrieval-Augmented Generation (RAG) pipelines for domain-specific use cases.
- Build AI agents, prompt orchestration frameworks, and multi-step reasoning workflows.
- Implement prompt engineering strategies, prompt templates, and context management.
- Develop guardrails, hallucination detection, moderation pipelines, and AI safety mechanisms.
- Build evaluation frameworks for model performance, prompt effectiveness, and business KPIs.
- Manage the end-to-end ML lifecycle from data collection to production deployment.
- Fine-tune foundation models using techniques such as LoRA, QLoRA, PEFT, and instruction tuning.
- Conduct supervised fine-tuning and experimentation on domain-specific datasets.
- Support advanced training approaches including transfer learning and model adaptation.
- Work with structured, unstructured, and multimodal datasets for training pipelines.
- Deploy scalable AI/ML solutions on AWS cloud infrastructure.
- Build and manage deployments using :
i. AWS SageMaker
ii. Amazon Bedrock
iii. AWS Lambda
iv. EKS / ECS
v. API Gateway
vi. S3 / Glue / Athena
- Build inference APIs and model-serving pipelines for low-latency production environments.
- Implement autoscaling and cost optimization strategies for model serving.
Model Optimization & Performance Engineering :
- Optimize model training and inference performance using :
i. DeepSpeed
ii. Quantization techniques
iii. Mixed precision training
iv. GPU optimization
- Work with GPU clusters and distributed training environments.
- Optimize vector search and embedding pipelines using vector databases.
- Build scalable retrieval pipelines using embeddings and semantic search.
- Work with vector databases such as :
i. Pinecone
ii. Weaviate
iii. Chroma
iv. FAISS
v. Milvus
- Design chunking, indexing, reranking, and retrieval strategies.
- Improve search relevance, latency, and retrieval quality.
- Establish end-to-end MLOps pipelines including CI/CD for ML workflows.
- Implement model versioning, experiment tracking, and monitoring.
- Work with tools such as MLflow, Weights & Biases, or equivalent.
- Set up drift monitoring, alerting, retraining workflows, and performance dashboards.
- Ensure compliance with responsible AI, governance, explainability, and auditability requirements.
- Perform exploratory data analysis (EDA) and feature research for model development.
- Build and maintain ETL/data pipelines for ML training and inference.
- Work with big data frameworks such as Spark or Flink.
- Conduct domain-focused research to improve model performance and business outcomes.
- Collaborate with product managers, UX teams, engineering teams, and stakeholders.
- Translate business requirements into scalable ML/AI solutions.
- Mentor junior engineers and guide technical best practices.
- Drive architecture decisions, code reviews, and deployment standards.
Required Skills & Qualifications :
- 5-12 years of experience in Machine Learning / AI Engineering.
- Minimum 3+ years of experience building and deploying production ML systems.
- 1+ years of hands-on experience with LLMs, RAG systems, and generative AI applications.
- Strong programming skills in Python.
- Hands-on experience with :
i. PyTorch
ii. Hugging Face Transformers
iii. LangChain / LlamaIndex
iv. OpenAI / Anthropic / open-source LLMs
- Strong AWS cloud expertise :
i. SageMaker
ii. Bedrock
iii. Lambda
iv. ECS/EKS
v. S3
- Experience with containerization and orchestration :
i. Docker
ii. Kubernetes
- Strong knowledge of CI/CD pipelines and ML monitoring.
- Experience with vector databases and embedding models.
- Strong understanding of ML system design and production architecture.
Did you find something suspicious?