HamburgerMenu
hirist

AI Infrastructure Engineer - Python/LLM/GPU

LanceTech Solutions
5 - 10 Years
rupee27-32 LPA
Multiple Locations

Posted on: 02/07/2026

Job Description

Title : AI Infrastructure Engineer (Python | LLM | GPU | AWS)

Location : Bangalore / Hyderabad / Pune / Chennai / Mumbai / Noida

Experience : 5+ Years in AI/ML Infra role, Python, Pytorch, Amazon Sagemaker, Multi-GPU

Overview :

We are looking for a Senior AI Infrastructure Engineer with hands-on experience building and optimizing AI infrastructure for Large Language Models (LLMs). The ideal candidate should have strong expertise in Python, PyTorch, AWS, GPU-based AI workloads, and production-scale ML systems.

Required Skills :

- 5+ years of experience in AI Infrastructure, ML Platform Engineering, MLOps, or Machine Learning Infrastructure.

- Strong programming experience in Python.

- Hands-on experience with PyTorch.

- Experience deploying, fine-tuning, or serving Large Language Models (LLMs).

- Experience building AI pipelines using Amazon SageMaker.

- Hands-on experience working with GPU-based AI workloads.

- Strong understanding of GPU architecture and performance optimization.

- Experience with distributed model training and multi-GPU environments.

- Good understanding of AWS services including : EC2 GPU Instances, SageMaker, S3, VPC, Networking.

- Experience designing scalable production Machine Learning systems.

- Good understanding of MLOps, model deployment, monitoring, and production lifecycle.

Key Responsibilities :

- Design, deploy, and manage Large Language Models (LLMs) on AWS infrastructure.

- Build, optimize, and maintain model training, fine-tuning, and inference pipelines using Amazon SageMaker.

- Optimize GPU-based AI workloads for training and inference performance.

- Design scalable ML infrastructure supporting distributed training and inference.

- Work with NVIDIA GPU environments (A100, H100) and AWS AI services including Trainium and Inferentia.

- Improve compute, networking, and storage performance for large-scale AI workloads.

- Benchmark model performance and optimize infrastructure for latency, throughput, and cost.

- Troubleshoot infrastructure bottlenecks across GPUs, distributed systems, and cloud environments.

- Collaborate with AI researchers, ML engineers, and platform teams to build production-ready AI systems.

- Implement infrastructure best practices for scalability, monitoring, reliability, and security.

Preferred Skills (Not Mandatory) :

- Experience working with NVIDIA A100/H100 GPUs.

- Exposure to AWS Trainium and Inferentia.

- Experience with Hugging Face Transformers and LLM deployment.

- Experience with distributed training frameworks.

- Experience with Kubernetes and Docker.

- Understanding of high-performance computing (HPC) or distributed systems.

- Exposure to model optimization techniques such as quantization, pruning, or knowledge distillation.

- Experience building high-throughput, low-latency AI inference systems.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Posted by

Hemalatha

TA at LanceTech Solutions

Last Active: NA as recruiter has posted this job through third party tool.

Job Views:  
105
Applications:  52
Recruiter Actions:  18

Posted in

AI/ML

Functional Area

Data Engineering

Job Code

1650569

Loading chat...