Posted on: 07/05/2026
Role : LLM infra & Optimization- AWS
Location : Chennai, Mumbai, Hyderabad, Pune, Noida, Gurgaon, Bangalore
Role Overview :
Own end-to-end deployment and optimization of large language models on AWS. Drive performance, scalability, and cost efficiency across training and inference workloads.
Key Responsibilities :
- Deploy and scale LLMs using Amazon SageMaker (training, fine-tuning, inference)
- Build and optimize ML pipelines for production workloads
- Improve model performance via GPU/CUDA-level tuning and infra optimization
- Design scalable architectures using AWS (EC2 GPU, S3, VPC, EFS)
- Implement distributed training and multi-GPU orchestration
- Benchmark performance and optimize cost vs latency
- Partner with internal teams/customers on deployment strategy and best practices
Core Requirements :
- 5+ years in ML infrastructure / model deployment / GPU computing
- Strong Python + experience with PyTorch / TensorFlow / JAX
- Hands-on experience deploying or fine-tuning LLMs in production
- Solid understanding of distributed training and inference systems
- Working knowledge of AWS core services (EC2, S3, networking)
Preferred :
- Experience with GPU optimization / CUDA
- Exposure to NVIDIA GPUs (A100/H100) or AWS Trainium/Inferentia
- Knowledge of model optimization (quantization, pruning, distillation)
- Familiarity with MLOps and production ML systems
Did you find something suspicious?