Posted on: 16/06/2026
Role Overview :
- As a Lead Site Reliability Engineer, you will serve as the technical anchor for our mission-critical infrastructure, ensuring high availability, scalability, and performance across our global cloud footprint.
- You will lead a high-performing team of engineers to bridge the gap between development and operations, fostering a culture of automation and proactive system health.
- By collaborating closely with product engineering teams and senior stakeholders, you will architect robust solutions that minimize downtime and optimize resource utilization.
- Your work directly impacts the end-user experience by ensuring seamless service delivery and maintaining the integrity of our production environments, ultimately driving business reliability and operational excellence.
Key Responsibilities :
Incident Response & Leadership :
- Lead and mentor a team of 5-6 SREs in daily operations and incident response activities.
- Act as the first responder to production alerts, rapidly assessing severity and initiating mitigation.
- Serve as Incident Commander during major incidents, leading bridge calls with clarity and decisiveness.
- Drive root cause isolation within 30 minutes for critical incidents whenever possible.
- Communicate effectively with engineering, product, and leadership teams during high-pressure situations.
- Maintain strong presence and ownership on incident bridges with confident decision-making.
Team Management & Operations :
- Oversee daily activities and coordinate closely with client leads and managers.
- Prepare weekly and monthly KPI reports on team performance and reliability metrics.
- Drive continuous improvement initiatives across the team.
Proactive Reliability Engineering :
- Identify patterns, trends, and signals to prevent incidents before they occur.
- Continuously improve alert quality, reduce noise, and increase signal fidelity.
- Partner with engineering teams to enhance system resilience and reliability.
Automation & Toil Reduction :
- Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows.
- Build and improve tooling across incident response, observability, and operations.
- Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value.
Platform & Systems Support :
- Troubleshoot across hybrid ecosystems including on-prem VMs (Linux & Windows; VMware), cloud platforms (AWS, GCP, Azure), and containerized environments (Kubernetes).
- Diagnose and resolve issues across networking, Kubernetes, CDN, and traffic management layers (Akamai, waiting rooms, etc.).
Required Skillset :
Core Engineering & Operations :
- Strong hands-on experience in incident management and triage in production environments.
- Proven ability to troubleshoot complex distributed systems under pressure.
- Solid understanding of Linux systems administration (performance tuning, networking, NTP, etc.).
- Prior experience as an Incident Commander or in similar leadership roles during outages.
Cloud & Infrastructure :
- Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2).
- Familiarity with GCP and/or Azure environments.
- Proven experience operating in multi-cloud and hybrid environments.
- Experience troubleshooting Kubernetes clusters (pods, ingress, configuration issues).
- Understanding of containerized application architectures.
DevOps & CI/CD :
- Strong knowledge of DevOps practices and CI/CD pipelines.
- Hands-on experience with Harness, GitHub, and/or GitLab.
Application & Technology Stack :
- Understanding of database connectivity and dependencies across Oracle, MariaDB, and MSSQL.
Networking :
- Strong foundational knowledge of TCP/IP, DNS, and HTTP(S).
- Experience with load balancing and network troubleshooting.
- Ability to diagnose connectivity issues between services and databases.
Preferred Qualifications :
- Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications.
- Familiarity with Akamai CDN and traffic management tools.
- Experience in high-volume, high-availability production environments.
- Track record of leading teams in fast-paced, incident-driven environments.
Working Hours :
- Shift 1 : 3 :00 AM - 12 :30 PM IST (5 :30 PM - 3 :00 AM EST)
- Shift 2 : 11 :00 AM - 8 :30 PM IST (1 :30 AM - 11 :00 AM EST)
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1645527