HamburgerMenu
hirist

Senior Site Reliability Engineer - Google Cloud Platform

Varite
7 - 10 Years
Multiple Locations

Posted on: 11/07/2026

Job Description

Role: Senior SRE.

Experience: 7 years.

Location: Bangalore/ Hyderabad.

Work Type: Contract - 1 year.

CTC: 25-30% hike on last drawn.

Interview: Virtual.

Mandate Skills:

- GCP (Google Cloud Platform) Must have (strong hands-on).

- GKE (Google Kubernetes Engine).

- Terraform (Infrastructure as Code).

- GitHub & GitHub Actions (CI/CD).

- SRE concepts (SLI, SLO, Monitoring, Incident Management, RCA).

- Cloud Networking & Security.

- Monitoring & Observability.

- REST APIs & Backend Automation.

Description:

Company Name: VARITE India Private Limited.

About The Client: Client is one of the worlds leading professional services firms and the fastest growing Big Four accounting firm.

Qualifications:

- 7+ years experience.

- Cloud Platform GCP (Mandatory).

- Strong expertise in designing and operating cloud-native architectures on GCP & AWS.

- Hands-on experience with core GCP services including Compute Engine, Networking, Access Management, Monitoring, Security & familiarity with Serverless architecture.

- Experience building secure, scalable, multi-region deployments.

- Knowledge of networking, security, and access control best practices.

- Familiarity with GCP observability tools.

- Familiarity with managing the Backup & recovery of cloud resources to maintain the desired RTO & RPO.

- Kubernetes & Container Orchestration (Core Skill).

- Deep understanding of Kubernetes architecture and control plane components.

- Strong expertise in managing container workloads at scale.

- Experience with workload scheduling, autoscaling, and resource optimization.

- Strong experience in creating custom network policies & managing the security posture at Cluster, Infrastructure & Application layers.

- Ability to troubleshoot cluster-level and application-level issues.

- Familiarity with network policies, RBAC, storage classes, and service mesh concepts.

- GKE.

- Hands-on experience with managed Kubernetes services (GKE).

- Expertise in cluster provisioning, upgrades, and node pool management.

- Experience with GKE security, workload identity, and cluster hardening.

- Ability to manage high availability and production-grade Kubernetes clusters.

- Infrastructure as Code Terraform.

- Strong experience in provisioning and managing infrastructure using Terraform.

- Ability to design and maintain reusable & secure Terraform self-service templates.

- Experience in state management, remote backends, and environment segregation.

- Integration of Terraform with CI/CD pipelines for automated deployments.

- Knowledge of infrastructure versioning and lifecycle management.

- Ability to integrate security scanning tools with IaC.

- Node.js & TypeScript.

- Proficiency in developing backend services and APIs using Node.js.

- Strong experience with TypeScript for scalable and maintainable codebases.

- Ability to build automation tools, services, and resource management utilities.

- Understanding of asynchronous programming, event-driven architecture, and REST APIs.

- Exposure to performance optimization and debugging of backend services.

- CI/CD & Version Control GitHub Ecosystem.

- Strong expertise in Git-based version control using GitHub.

- Experience in building and maintaining CI/CD pipelines using GitHub Actions or similar tools.

- Knowledge of branching strategies, code reviews, and release workflows.

- Ability to automate build, test, and deployment processes for cloud-native applications.

- Helm.

- Hands-on experience with Helm for packaging and deploying Kubernetes applications.

- Ability to create and maintain custom Helm charts.

- Experience in managing application lifecycle, versioning, and environment-specific configurations.

- Understanding of templating and release management using Helm.

- SRE Practices.

- Strong understanding of SRE principles including SLIs, SLOs, and error budgets.

- Experience in designing highly reliable and resilient systems.

- Expertise in monitoring, logging, and alerting strategies for distributed systems.

- Ability to drive capacity planning, resource allocation, and cost optimization.

- Hands-on experience in incident management, root cause analysis, and postmortems.

- Experience in improving system performance, availability, and scalability.

- Multi-Cloud Exposure AWS/Azure (Good to Have).

- Basic exposure to AWS services.

- Understanding of multi-cloud environments.

- Understanding of implementing the security posture on the cloud.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...