Posted on: 19/08/2026
Experience: 510 Years
Location: Bengaluru
Work Mode: On-site
Role Overview:
We are looking for a Senior GCP DevOps Engineer with strong hands-on experience in GCP, Kubernetes/GKE, Terraform, CI/CD, networking, and observability. The candidate should have experience managing production-grade cloud infrastructure and exposure to AI/ML workloads, including GPU-enabled infrastructure.
Key Responsibilities:
- Build, manage, and optimize production-grade infrastructure on GCP.
- Manage Kubernetes/GKE clusters, including deployments, autoscaling, resource management, and production troubleshooting.
- Implement Infrastructure as Code using Terraform.
- Design and maintain CI/CD pipelines using Jenkins, GitLab CI, GitHub Actions, Cloud Build, or ArgoCD.
- Implement monitoring and observability using Cloud Monitoring, Cloud Logging, Prometheus, and Grafana.
- Troubleshoot production issues related to Kubernetes, networking, resource utilization, and application performance.
- Design and manage cloud networking, including VPCs, subnets, routing, firewall rules, DNS, and load balancing.
- Support AI/ML workloads, including GPU-enabled infrastructure and model-serving environments.
- Work closely with ML and backend teams to build scalable and reliable infrastructure.
- Drive cloud cost optimization and infrastructure reliability.
Required Skills:
- Strong hands-on experience with GCP and GKE.
- Strong Kubernetes production experience.
- Good hands-on experience with cloud networking and troubleshooting.
- Strong understanding of VPC, subnets, routing, firewall rules, DNS, and network security.
- Strong experience with load balancers, traffic routing, and high-availability configurations.
- Strong Terraform/IaC experience.
- Experience with CI/CD and GitOps.
- Experience with Prometheus, Grafana, Cloud Monitoring, and Cloud Logging.
- Experience with GCP IAM, autoscaling, and infrastructure optimization.
- Experience supporting AI/ML workloads or GPU infrastructure is preferred.
- Exposure to AWS services such as EKS, EC2, IAM, and VPC.
- Strong troubleshooting and problem-solving skills.
Preferred Skills:
- Experience with GPU node management and ML workload scheduling.
- Exposure to Triton, TorchServe, or vLLM.
- Experience with cloud cost optimization.
- Strong scripting skills in Python or Bash.
- Experience working closely with ML and backend engineering teams.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1664464