Posted on: 22/09/2026
Job Description:
Immediate joiners or less than 30 days notice period.
Job Summary:
We are looking for an AI Infrastructure Engineer with 12+ years of total experience and 2 - 4 years of hands-on experience in AI/ML infrastructure or MLOps. The role focuses on designing, scaling, and optimizing AI platforms, cloud infrastructure, automation, and reliable ML/LLM deployment environments.
Key Responsibilities:
- Design, build, and maintain scalable infrastructure, pipelines, and environments for AI/ML model training and deployment.
- Partner with Infrastructure Operations teams to identify operational bottlenecks and implement AI-driven automation.
- Optimize GPU/CPU workloads, storage, and compute clusters for performance and cost efficiency.
- Translate business and operational requirements into AI infrastructure initiatives.
- Measure infrastructure improvements through uptime, latency, cloud cost, and productivity metrics.
- Establish MLOps and LLMOps practices including CI/CD, monitoring, drift detection, and model governance.
- Implement observability and telemetry for AI models and underlying infrastructure.
- Develop reliable, scalable, and resilient AI infrastructure solutions.
Mandatory Skills:
- 12+ years of experience in Software or Infrastructure Engineering.
- 2 - 4 years of experience in AI/ML infrastructure or MLOps.
- Strong experience in AWS or Azure cloud architecture.
- Hands-on experience with cloud networking, IAM, security, and infrastructure architecture.
- Strong experience with Terraform/IaC and infrastructure automation.
- Hands-on experience with Kubernetes and Docker.
- Proficiency in Python or Go automation and strong Linux/shell scripting skills.
- Experience with observability and ITSM integration.
- Experience integrating AI/LLM APIs into IT operations.
- Exposure to RAG and AI-driven automation solutions.
AI/ML & MLOps Skills:
- Experience deploying AI/ML frameworks such as PyTorch and TensorFlow.
- Experience with vector databases.
- Hands-on experience with MLOps tools such as Kubeflow, MLflow, Ray, or Triton Inference Server.
- Understanding of LLM infrastructure, model deployment, and AI platform scalability.
Preferred Skills:
- Experience with LLM infrastructure and fine-tuning environments.
- Experience building and scaling RAG pipelines.
- AWS, Azure, or GCP Architecture certification.
- Kubernetes certifications such as CKA or CKAD.
- Experience working with global distributed teams.
Key Competencies:
- Strong business and cost-management mindset.
- Excellent problem-solving and automation skills.
- Strong communication and cross-functional collaboration.
- Ability to connect infrastructure decisions with business outcomes, SLAs, scalability, and cost efficiency.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1673571