HamburgerMenu
hirist

OpsTree - Senior Site Reliability Engineer - Google Cloud Platform

OpsTree Solutions
7 - 8 Years
Bangalore

Posted on: 11/09/2026

Job Description

Experience : 7 - 8 Years

Employment Type : Full-Time

Role : Senior Site Reliability Engineer (SRE) - GCP

Job Summary :

We are looking for an experienced Senior Site Reliability Engineer (SRE) with 7 - 8 years of experience in designing, implementing, and maintaining highly available, scalable, and reliable cloud infrastructure on Google Cloud Platform (GCP).

The ideal candidate should have strong hands-on experience with GCP, Kubernetes, Linux, Infrastructure as Code, CI/CD, observability, automation, and production operations. The role will involve improving system reliability, performance, scalability, and operational efficiency while working closely with development and platform teams.

Key Responsibilities :

- Design, deploy, and manage highly available and scalable infrastructure on GCP.

- Own production infrastructure, reliability, availability, performance, and operational excellence.

- Build and maintain GKE/Kubernetes clusters and containerized workloads.

- Implement and manage CI/CD pipelines for application and infrastructure deployments.

- Automate repetitive operational tasks using Python, Bash, or similar scripting languages.

- Develop and maintain Infrastructure as Code using Terraform.

- Implement monitoring, logging, alerting, and observability using Google Cloud Monitoring, Cloud Logging, and other tools.

- Define and monitor SLIs, SLOs, and SLAs for critical services.

- Participate in production incident management, troubleshooting, root-cause analysis, and post-incident reviews.

- Perform capacity planning, performance tuning, and scalability assessments.

- Implement high-availability, disaster recovery, backup, and business-continuity strategies.

- Identify reliability risks and proactively implement solutions to reduce system downtime.

- Work with development teams to improve application reliability and deployment processes.

- Establish and improve operational runbooks, automation, and self-healing mechanisms.

- Support security and compliance requirements across cloud infrastructure.

- Participate in on-call/production support activities when required.

Required Technical Skills :

GCP :

- Strong hands-on experience with Google Cloud Platform.

- Experience with GKE, Compute Engine, VPC, IAM, Cloud Load Balancing, Cloud Storage, Cloud SQL, Pub/Sub, Cloud Monitoring & Logging, Secret Manager.

- Good understanding of GCP networking, IAM, security, and cost optimization.

Kubernetes & Containers :

- Strong hands-on experience with Kubernetes and GKE.

- Docker/containerization.

- Kubernetes networking, deployments, services, ingress, configmaps, secrets, RBAC, and troubleshooting.

- Experience with Helm and Kubernetes operators is preferred.

Infrastructure as Code :

- Strong experience with Terraform.

- Experience designing reusable Terraform modules and managing infrastructure through Git-based workflows.

CI/CD & DevOps :

- Strong understanding of CI/CD principles.

- Hands-on experience with tools such as GitLab CI/CD, Jenkins, GitHub Actions, ArgoCD.

- Experience implementing automated deployment and rollback strategies.

Observability & Reliability :

- Experience with monitoring, logging, tracing, and alerting.

- Hands-on experience with Prometheus, Grafana, OpenTelemetry or equivalent tools.

- Strong understanding of SLIs, SLOs, SLAs, Error Budgets, Incident Management, RCA.

Linux & Networking :

- Strong Linux administration and troubleshooting skills.

- Good understanding of TCP/IP, DNS, HTTP/HTTPS, Load Balancing, SSL/TLS, Network troubleshooting, Firewall and security concepts.

Scripting & Automation :

- Strong scripting experience in Python and/or Bash.

- Ability to build automation tools and operational scripts.

Good to Have :

- GCP Professional Cloud DevOps Engineer certification.

- Experience with GitOps and ArgoCD.

- Experience with service mesh such as Istio.

- Experience with SRE frameworks and practices.

- Experience with chaos engineering and resilience testing.

- Experience with distributed systems and microservices architecture.

- Experience with FinOps/cloud cost optimization.

- Exposure to security best practices and DevSecOps.

- Experience working in 24x7 production environments.

Required Soft Skills :

- Strong analytical and problem-solving skills.

- Excellent troubleshooting and debugging ability.

- Good communication and documentation skills.

- Ability to work independently and take ownership of production systems.

- Strong incident management and decision-making skills.

- Ability to collaborate effectively with Development, Security, Product, and Platform teams.

Education :

Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field is preferred.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...