Posted on: 04/07/2026
Note : This is a direct opening with one of our Esteemed Clients.
We are seeking a highly skilled Site Reliability Engineer (SRE) with deep expertise in Kubernetes (GKE) and high-scale data platforms. This role is critical for a candidate who is not just a user of tools, but an expert in managing the lifecycle of ELK clusters, Redis Sentinels, and Apache Kafka. You will be responsible for automating our infrastructure, onboarding new applications, and conducting Proof of Concepts (POCs) for multi-tier application automation.
Key Responsibilities :
- Kubernetes Mastery : Design, manage, and optimize GKE (Google Kubernetes Engine) clusters. Act as the subject matter expert for all K8s-related tasks, including resource scaling, networking, and security.
- ELK Stack Administration : Full ownership of the ELK (Elasticsearch, Logstash, Kibana) infrastructure. This includes :
1. Onboarding new applications and log sources.
2. Managing User Access (RBAC) and security roles.
3. Creating advanced Kibana dashboards and alerting systems.
- Data Tier Management :
1. Redis : Deploy and manage Redis Sentinel for high availability; handle instance creation and performance tuning.
2. Apache Kafka : Manage Kafka clusters, including topic creation, replication factor management, and partition balancing.
- Automation & POCs : Drive innovation by performing POCs for multi-tier applications. Build custom automation to reduce manual toil across the entire stack.
- Cloud Infrastructure : Manage GCP resources using Infrastructure as Code (Terraform/Ansible) with a focus on cost-efficiency and 99.99% availability.
Required Technical Skills :
- Orchestration : Expert-level knowledge of Kubernetes (GKE) and Docker.
- Observability : Deep experience in ELK Stack administration (not just searching logs, but managing the cluster health and user permissions).
- Messaging & Caching : Hands-on experience managing Apache Kafka (Topics/Replication) and Redis (Sentinel/Clustering).
- Automation : Proficiency in Python or Java for building automation tools and conducting complex technical POCs.
- CI/CD : Experience with ArgoCD, Jenkins, or GitLab CI/CD for automated application delivery.
- Cloud : Strong knowledge of GCP (VPC, IAM, GKE, Cloud Storage).
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1651441