Posted on: 24/06/2026
Lead Platform Engineer (Observability Focus)
About The Role :
We are looking for a Senior Platform Engineer with deep expertise in observability, cloud-native infrastructure, and large-scale distributed systems. This role is highly hands-on and focuses on designing, building, and operating reliable, observable, and scalable platforms running on Kubernetes, with a strong preference for AWS, and having GCP knowledge would be an edge.
Key Responsibilities :
Reliability & Operations :
- Design, implement, and maintain highly available and resilient systems in Kubernetes-based environments.
- Define and enforce SLOs, SLIs, and error budgets.
- Lead incident response, RCA, and postmortems.
- Drive reliability improvements through automation.
Observability (Core Focus) :
- Architect and operate observability platforms for metrics, logging, tracing, and alerting.
- Work with Prometheus, Alertmanager, Grafana, Splunk, Cribl, Datadog.
- Establish actionable alerting standards.
Cloud & Platform Engineering :
- Build and manage infrastructure on AWS.
- Operate Kubernetes clusters (EKS preferred).
- Deploy services using Helm, ArgoCD and Argo rollout.
- Manage containerized workloads using Docker and containerd.
Automation & Tooling :
- Strong Python skills with emphasis on reliability, automation, and observability tooling.
- Develop automation and tooling using Python.
- Create internal reliability and monitoring tools.
- Integrate CI/CD pipelines with observability and reliability checks.
Collaboration & Leadership :
- Mentor junior engineers.
- Influence architecture decisions.
- Collaborate across engineering teams.
Required Qualifications :
- 6+ years of relevant experience in SRE, DevOps, or Platform Engineering.
- Strong Python skills with experience building production-grade automation and tooling.
- Strong programming experience in Python.
- Production experience with Kubernetes.
- Strong observability fundamentals.
- Experience with Helm, ArgoCD, Argo Rollout and Docker.
- Experience with AWS cloud.
- Strong Linux and networking fundamentals.
- Familiarity with the SDLC.
Preferred Qualifications :
- Experience with OpenTelemetry and Observability tools.
- Experience with Kubernetes package manager (helm) and deployment (ArgoCD / Argo Rollout).
- Multi-cluster or multi-region Kubernetes experience.
- Service mesh (Istio) and API Gateway (Kong) experience.
- Infrastructure-as-Code (Terraform preferred).
- Cloud cost optimization experience.
Technology Stack :
- Programming & Automation : Python (strong proficiency, production-grade tooling and automation).
- Containerization & Orchestration : Docker, AWS EKS.
- Packaging & Deployment : Helm, ArgoCD, Argo Rollout.
- Observability & Monitoring : Prometheus, Alertmanager, Grafana, OpenTelemetry, Datadog, Splunk, Cribl, Edge Collectors, AWS Cloud Watch.
- Platforms : AWS (primary and preferred), Google Cloud Platform (good to have).
- CI/CD & DevOps : Git-based CI/CD pipelines, release automation, reliability checks.
- Infrastructure as Code : Terraform (preferred).
- Operating Systems & Networking : Linux, TCP/IP, DNS, load balancing.
Project Details / What Youll Work On :
- Build and operate a centralized observability platform for metrics, logs, traces, and alerting across Kubernetes workloads using above mentioned tooling for services running in AWS, on-prem, and GCP (good to have).
- Contribute to the o11y center-of-excellence guiding teams to create their SLOs, SLIs, reduce MTTR and collaborate with developers to move toward and Observability Driven Development mindset.
- Support observability for services running on our Kubernetes ecosystem.
- Develop Python-based automation and tooling for observability, SLO reporting, incident response, and operational workflows.
- Lead incident response for production issues, conduct blameless postmortems, and drive long-term reliability improvements.
- Optimize platform scalability, performance, and cloud cost efficiency with a strong focus on GCP and AWS.
- Act as a technical leader, influencing architecture and mentoring teams on reliability and observability best practices.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1647866