HamburgerMenu
hirist

Opstree Solutions - Senior GCP DevOps Engineer - Python/Kubernetes

OpsTree Solutions
5 - 10 Years
Bangalore

Posted on: 19/08/2026

Job Description

Experience: 510 Years

Location: Bengaluru

Work Mode: On-site

Role Overview:

We are looking for a Senior GCP DevOps Engineer with strong hands-on experience in GCP, Kubernetes/GKE, Terraform, CI/CD, networking, and observability. The candidate should have experience managing production-grade cloud infrastructure and exposure to AI/ML workloads, including GPU-enabled infrastructure.

Key Responsibilities:

- Build, manage, and optimize production-grade infrastructure on GCP.

- Manage Kubernetes/GKE clusters, including deployments, autoscaling, resource management, and production troubleshooting.

- Implement Infrastructure as Code using Terraform.

- Design and maintain CI/CD pipelines using Jenkins, GitLab CI, GitHub Actions, Cloud Build, or ArgoCD.

- Implement monitoring and observability using Cloud Monitoring, Cloud Logging, Prometheus, and Grafana.

- Troubleshoot production issues related to Kubernetes, networking, resource utilization, and application performance.

- Design and manage cloud networking, including VPCs, subnets, routing, firewall rules, DNS, and load balancing.

- Support AI/ML workloads, including GPU-enabled infrastructure and model-serving environments.

- Work closely with ML and backend teams to build scalable and reliable infrastructure.

- Drive cloud cost optimization and infrastructure reliability.

Required Skills:

- Strong hands-on experience with GCP and GKE.

- Strong Kubernetes production experience.

- Good hands-on experience with cloud networking and troubleshooting.

- Strong understanding of VPC, subnets, routing, firewall rules, DNS, and network security.

- Strong experience with load balancers, traffic routing, and high-availability configurations.

- Strong Terraform/IaC experience.

- Experience with CI/CD and GitOps.

- Experience with Prometheus, Grafana, Cloud Monitoring, and Cloud Logging.

- Experience with GCP IAM, autoscaling, and infrastructure optimization.

- Experience supporting AI/ML workloads or GPU infrastructure is preferred.

- Exposure to AWS services such as EKS, EC2, IAM, and VPC.

- Strong troubleshooting and problem-solving skills.

Preferred Skills:

- Experience with GPU node management and ML workload scheduling.

- Exposure to Triton, TorchServe, or vLLM.

- Experience with cloud cost optimization.

- Strong scripting skills in Python or Bash.

- Experience working closely with ML and backend engineering teams.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...