HamburgerMenu
hirist

AI Infrastructure Engineer

Okda Solutions
7 - 10 Years
Pune

Posted on: 22/09/2026

Job Description

Job Description:

Immediate joiners or less than 30 days notice period.

Job Summary:

We are looking for an AI Infrastructure Engineer with 12+ years of total experience and 2 - 4 years of hands-on experience in AI/ML infrastructure or MLOps. The role focuses on designing, scaling, and optimizing AI platforms, cloud infrastructure, automation, and reliable ML/LLM deployment environments.

Key Responsibilities:

- Design, build, and maintain scalable infrastructure, pipelines, and environments for AI/ML model training and deployment.

- Partner with Infrastructure Operations teams to identify operational bottlenecks and implement AI-driven automation.

- Optimize GPU/CPU workloads, storage, and compute clusters for performance and cost efficiency.

- Translate business and operational requirements into AI infrastructure initiatives.

- Measure infrastructure improvements through uptime, latency, cloud cost, and productivity metrics.

- Establish MLOps and LLMOps practices including CI/CD, monitoring, drift detection, and model governance.

- Implement observability and telemetry for AI models and underlying infrastructure.

- Develop reliable, scalable, and resilient AI infrastructure solutions.

Mandatory Skills:

- 12+ years of experience in Software or Infrastructure Engineering.

- 2 - 4 years of experience in AI/ML infrastructure or MLOps.

- Strong experience in AWS or Azure cloud architecture.

- Hands-on experience with cloud networking, IAM, security, and infrastructure architecture.

- Strong experience with Terraform/IaC and infrastructure automation.

- Hands-on experience with Kubernetes and Docker.

- Proficiency in Python or Go automation and strong Linux/shell scripting skills.

- Experience with observability and ITSM integration.

- Experience integrating AI/LLM APIs into IT operations.

- Exposure to RAG and AI-driven automation solutions.

AI/ML & MLOps Skills:

- Experience deploying AI/ML frameworks such as PyTorch and TensorFlow.

- Experience with vector databases.

- Hands-on experience with MLOps tools such as Kubeflow, MLflow, Ray, or Triton Inference Server.

- Understanding of LLM infrastructure, model deployment, and AI platform scalability.

Preferred Skills:

- Experience with LLM infrastructure and fine-tuning environments.

- Experience building and scaling RAG pipelines.

- AWS, Azure, or GCP Architecture certification.

- Kubernetes certifications such as CKA or CKAD.

- Experience working with global distributed teams.

Key Competencies:

- Strong business and cost-management mindset.

- Excellent problem-solving and automation skills.

- Strong communication and cross-functional collaboration.

- Ability to connect infrastructure decisions with business outcomes, SLAs, scalability, and cost efficiency.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...