Posted on: 02/07/2026
Title : AI Infrastructure Engineer (Python | LLM | GPU | AWS)
Location : Bangalore / Hyderabad / Pune / Chennai / Mumbai / Noida
Experience : 5+ Years in AI/ML Infra role, Python, Pytorch, Amazon Sagemaker, Multi-GPU
Overview :
We are looking for a Senior AI Infrastructure Engineer with hands-on experience building and optimizing AI infrastructure for Large Language Models (LLMs). The ideal candidate should have strong expertise in Python, PyTorch, AWS, GPU-based AI workloads, and production-scale ML systems.
Required Skills :
- 5+ years of experience in AI Infrastructure, ML Platform Engineering, MLOps, or Machine Learning Infrastructure.
- Strong programming experience in Python.
- Hands-on experience with PyTorch.
- Experience deploying, fine-tuning, or serving Large Language Models (LLMs).
- Experience building AI pipelines using Amazon SageMaker.
- Hands-on experience working with GPU-based AI workloads.
- Strong understanding of GPU architecture and performance optimization.
- Experience with distributed model training and multi-GPU environments.
- Good understanding of AWS services including : EC2 GPU Instances, SageMaker, S3, VPC, Networking.
- Experience designing scalable production Machine Learning systems.
- Good understanding of MLOps, model deployment, monitoring, and production lifecycle.
Key Responsibilities :
- Design, deploy, and manage Large Language Models (LLMs) on AWS infrastructure.
- Build, optimize, and maintain model training, fine-tuning, and inference pipelines using Amazon SageMaker.
- Optimize GPU-based AI workloads for training and inference performance.
- Design scalable ML infrastructure supporting distributed training and inference.
- Work with NVIDIA GPU environments (A100, H100) and AWS AI services including Trainium and Inferentia.
- Improve compute, networking, and storage performance for large-scale AI workloads.
- Benchmark model performance and optimize infrastructure for latency, throughput, and cost.
- Troubleshoot infrastructure bottlenecks across GPUs, distributed systems, and cloud environments.
- Collaborate with AI researchers, ML engineers, and platform teams to build production-ready AI systems.
- Implement infrastructure best practices for scalability, monitoring, reliability, and security.
Preferred Skills (Not Mandatory) :
- Experience working with NVIDIA A100/H100 GPUs.
- Exposure to AWS Trainium and Inferentia.
- Experience with Hugging Face Transformers and LLM deployment.
- Experience with distributed training frameworks.
- Experience with Kubernetes and Docker.
- Understanding of high-performance computing (HPC) or distributed systems.
- Exposure to model optimization techniques such as quantization, pruning, or knowledge distillation.
- Experience building high-throughput, low-latency AI inference systems.
Did you find something suspicious?