Posted on: 17/08/2026
Role Overview :
As a Senior AI Platform & MLOps Engineer, you will be the backbone of our machine learning infrastructure, bridging the gap between cutting-edge AI research and production-grade reliability. You will spend your days architecting scalable deployment pipelines, optimizing high-performance inference services, and ensuring our AI models operate with maximum efficiency on AWS. Working closely with Data Scientists, Product Managers, and DevOps teams, you will play a pivotal role in accelerating our time-to-market for AI features. Your work directly impacts the end-user experience by ensuring that our intelligent systems are not only fast and accurate but also resilient and highly available at scale.
Key Responsibilities :
- Engineer and maintain robust CI/CD pipelines to automate the end-to-end lifecycle of machine learning models, ensuring seamless transitions from development to production.
- Architect and manage high-performance inference services on EKS and Kubernetes, optimizing resource utilization to support heavy real-time traffic.
- Implement Infrastructure as Code (IaC) using Terraform to ensure environment consistency, scalability, and rapid provisioning across our cloud footprint.
- Integrate advanced monitoring and observability frameworks using OpenTelemetry to proactively identify bottlenecks and ensure system reliability.
- Leverage AWS Bedrock and related generative AI services to build scalable, secure, and cost-effective AI platforms that meet evolving business requirements.
- Collaborate with cross-functional engineering teams to establish best practices for model versioning, automated testing, and production monitoring.
Required Skillset :
- Demonstrated expertise in managing containerized workloads on Kubernetes and EKS, with a deep understanding of cluster orchestration and networking.
- Proven ability to design and maintain production-grade CI/CD workflows that support rapid iteration cycles for AI/ML teams.
- Strong proficiency in Infrastructure as Code (IaC) using Terraform, with a focus on building modular, reusable, and secure cloud infrastructure.
- Hands-on experience in optimizing model inference performance and implementing observability solutions using OpenTelemetry to maintain high system uptime.
- Ability to effectively communicate complex technical architectures to stakeholders and mentor junior engineers to foster a culture of technical excellence.
- Strong problem-solving mindset with the adaptability to thrive in a fast-paced, hybrid work environment based in Bangalore.
- A degree in Computer Science, Engineering, or a related quantitative field, complemented by 4 - 9 years of relevant experience in platform or MLOps engineering.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1663712