Posted on: 10/09/2026
Experience:
10+ years in Cloud Infrastructure, DevOps, SRE, or Platform Engineering, with strong hands-on experience in Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE).
Role Overview:
We are looking for an experienced GCP DevOps Architect to design, implement, and operate highly scalable, secure, resilient, and cost-efficient cloud platforms on GCP.
The ideal candidate should have strong expertise in GCP, GKE, Kubernetes, DevOps, CI/CD, Infrastructure as Code, observability, security, and high-availability architecture, with experience building platforms capable of handling 5 - 10 million users / high-volume production traffic.
The candidate will be responsible for defining cloud architecture, DevOps standards, deployment strategies, reliability engineering practices, and platform automation.
Key Responsibilities:
GCP & Cloud Architecture:
- Design highly available and scalable cloud architectures on GCP.
- Architect platforms capable of supporting 5 - 10 million users and high-volume traffic.
- Design multi-region / multi-zone architectures for high availability and disaster recovery.
- Select and optimize appropriate GCP services based on scalability, reliability, performance, and cost.
- Design networking architecture including VPC, subnets, Cloud Load Balancing, Cloud NAT, Cloud DNS, Private Service Connect, firewall policies, and hybrid connectivity.
- Define cloud architecture standards, reference architectures, and engineering best practices.
GKE & Kubernetes:
- Strong hands-on experience designing and managing GKE production clusters.
- Design scalable GKE architectures including regional and private GKE clusters, node pools, workload isolation, cluster autoscaling, HPA/VPA, Workload Identity, Ingress/Gateway architecture, network policies, and pod security.
- Optimize Kubernetes workloads for performance, availability, and cost.
- Define Kubernetes deployment, upgrade, backup, and disaster recovery strategies.
- Troubleshoot complex production issues involving Kubernetes, networking, compute, and application performance.
DevOps & CI/CD:
- Design and implement enterprise-grade CI/CD pipelines.
- Strong experience with tools such as GitHub/GitLab, Jenkins, Cloud Build, Argo CD, and/or other CI/CD platforms.
- Implement GitOps-based deployment strategies where appropriate.
- Automate build, test, security scanning, deployment, rollback, and release processes.
- Implement progressive delivery strategies such as Blue/Green, Canary, and Rolling deployments.
- Establish DevOps standards across development and operations teams.
Infrastructure as Code:
- Strong hands-on experience with Terraform.
- Build reusable Terraform modules and infrastructure frameworks.
- Automate provisioning and configuration of GCP infrastructure.
- Implement infrastructure versioning, state management, policy controls, and automated validation.
Scalability & Reliability:
- Architect systems for millions of users and high concurrent traffic.
- Design autoscaling strategies across compute, Kubernetes, databases, and networking layers.
- Implement SRE practices including SLIs/SLOs/SLAs, error budgets, capacity planning, reliability engineering, incident management, and performance engineering.
- Conduct architecture reviews, scalability assessments, and production readiness reviews.
- Design fault-tolerant systems with appropriate RTO/RPO targets.
Monitoring & Observability:
- Design comprehensive monitoring and observability solutions using Google Cloud Operations Suite, Prometheus, Grafana, OpenTelemetry, and Datadog.
- Implement infrastructure, application, Kubernetes, and business-level monitoring.
- Establish centralized logging, metrics, tracing, alerting, and dashboards.
- Analyze production performance and identify bottlenecks.
Security:
- Implement GCP security best practices across infrastructure and Kubernetes.
- Strong understanding of IAM, Service Accounts, Workload Identity, Secret Manager, KMS, VPC Service Controls, Organization Policies, Security Command Center, and container/image security.
- Implement least-privilege access and secure CI/CD pipelines.
- Integrate vulnerability scanning and security controls into the DevOps lifecycle.
Cost Optimization:
- Monitor and optimize GCP infrastructure costs.
- Optimize GKE compute, node pools, autoscaling, storage, networking, and logging costs.
- Establish cloud FinOps practices and cost governance.
- Identify opportunities for capacity optimization without compromising reliability.
Required Technical Skills:
- Must Have: Strong GCP expertise, Strong GKE / Kubernetes expertise, Strong Terraform / IaC, Advanced CI/CD and DevOps, GCP networking and security, Linux and container technologies, Production experience with highly scalable systems, Experience supporting 5 - 10 million users or equivalent high-volume traffic, High Availability and Disaster Recovery architecture, Monitoring, logging, and observability, Strong troubleshooting and incident-management skills.
- Good to Have: Google Cloud Professional Cloud Architect certification, Google Cloud Professional Cloud DevOps Engineer certification, Service mesh experience such as Istio, Argo CD / GitOps, Prometheus / Grafana / OpenTelemetry, Anthos / GKE Enterprise, FinOps experience, Experience with Kafka, Redis, PostgreSQL/MySQL, NoSQL platforms, Experience with API Gateway / Apigee, Experience with microservices architecture.
Leadership & Soft Skills:
- Strong architectural and problem-solving capabilities.
- Ability to communicate complex cloud architecture to both technical and non-technical stakeholders.
- Experience mentoring DevOps, SRE, and cloud engineering teams.
- Ability to lead architecture decisions and establish engineering standards.
- Strong ownership of production reliability and operational excellence.
- Ability to work effectively with application, security, data, and infrastructure teams.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1670303