HamburgerMenu
hirist

Sony - Lead Platform Engineer - Observability Services

Sony
8 - 12 Years
Bangalore

Posted on: 24/06/2026

Job Description

Lead Platform Engineer (Observability Focus)

About The Role :

We are looking for a Senior Platform Engineer with deep expertise in observability, cloud-native infrastructure, and large-scale distributed systems. This role is highly hands-on and focuses on designing, building, and operating reliable, observable, and scalable platforms running on Kubernetes, with a strong preference for AWS, and having GCP knowledge would be an edge.

Key Responsibilities :

Reliability & Operations :

- Design, implement, and maintain highly available and resilient systems in Kubernetes-based environments.

- Define and enforce SLOs, SLIs, and error budgets.

- Lead incident response, RCA, and postmortems.

- Drive reliability improvements through automation.

Observability (Core Focus) :

- Architect and operate observability platforms for metrics, logging, tracing, and alerting.

- Work with Prometheus, Alertmanager, Grafana, Splunk, Cribl, Datadog.

- Establish actionable alerting standards.

Cloud & Platform Engineering :

- Build and manage infrastructure on AWS.

- Operate Kubernetes clusters (EKS preferred).

- Deploy services using Helm, ArgoCD and Argo rollout.

- Manage containerized workloads using Docker and containerd.

Automation & Tooling :

- Strong Python skills with emphasis on reliability, automation, and observability tooling.

- Develop automation and tooling using Python.

- Create internal reliability and monitoring tools.

- Integrate CI/CD pipelines with observability and reliability checks.

Collaboration & Leadership :

- Mentor junior engineers.

- Influence architecture decisions.

- Collaborate across engineering teams.

Required Qualifications :

- 6+ years of relevant experience in SRE, DevOps, or Platform Engineering.

- Strong Python skills with experience building production-grade automation and tooling.

- Strong programming experience in Python.

- Production experience with Kubernetes.

- Strong observability fundamentals.

- Experience with Helm, ArgoCD, Argo Rollout and Docker.

- Experience with AWS cloud.

- Strong Linux and networking fundamentals.

- Familiarity with the SDLC.

Preferred Qualifications :

- Experience with OpenTelemetry and Observability tools.

- Experience with Kubernetes package manager (helm) and deployment (ArgoCD / Argo Rollout).

- Multi-cluster or multi-region Kubernetes experience.

- Service mesh (Istio) and API Gateway (Kong) experience.

- Infrastructure-as-Code (Terraform preferred).

- Cloud cost optimization experience.

Technology Stack :

- Programming & Automation : Python (strong proficiency, production-grade tooling and automation).

- Containerization & Orchestration : Docker, AWS EKS.

- Packaging & Deployment : Helm, ArgoCD, Argo Rollout.

- Observability & Monitoring : Prometheus, Alertmanager, Grafana, OpenTelemetry, Datadog, Splunk, Cribl, Edge Collectors, AWS Cloud Watch.

- Platforms : AWS (primary and preferred), Google Cloud Platform (good to have).

- CI/CD & DevOps : Git-based CI/CD pipelines, release automation, reliability checks.

- Infrastructure as Code : Terraform (preferred).

- Operating Systems & Networking : Linux, TCP/IP, DNS, load balancing.

Project Details / What Youll Work On :

- Build and operate a centralized observability platform for metrics, logs, traces, and alerting across Kubernetes workloads using above mentioned tooling for services running in AWS, on-prem, and GCP (good to have).

- Contribute to the o11y center-of-excellence guiding teams to create their SLOs, SLIs, reduce MTTR and collaborate with developers to move toward and Observability Driven Development mindset.

- Support observability for services running on our Kubernetes ecosystem.

- Develop Python-based automation and tooling for observability, SLO reporting, incident response, and operational workflows.

- Lead incident response for production issues, conduct blameless postmortems, and drive long-term reliability improvements.

- Optimize platform scalability, performance, and cloud cost efficiency with a strong focus on GCP and AWS.

- Act as a technical leader, influencing architecture and mentoring teams on reliability and observability best practices.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...