Posted on: 27/06/2026
Job Description :
Incident Response & Leadership :
- Lead and mentor a team of 5-6 SREs in daily operations and incident response activities
- Act as the first responder to production alerts, rapidly assessing severity and initiating mitigation
- Serve as Incident Commander during major incidents, leading bridge calls with clarity and decisiveness
- Drive root cause isolation within 30 minutes for critical incidents whenever possible
- Communicate effectively with engineering, product, and leadership teams during high-pressure situations
- Maintain strong presence and ownership on incident bridges with confident decision-making
Team Management & Operations :
- Oversee daily activities and coordinate closely with client leads and managers
- Prepare weekly and monthly KPI reports on team performance and reliability metrics
- Drive continuous improvement initiatives across the team
Proactive Reliability Engineering :
- Identify patterns, trends, and signals to prevent incidents before they occur
- Continuously improve alert quality, reduce noise, and increase signal fidelity
- Partner with engineering teams to enhance system resilience and reliability
Automation & Toil Reduction :
- Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows
- Build and improve tooling across incident response, observability, and operations
- Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value
Platform & Systems Support :
- Troubleshoot across hybrid ecosystems including on-prem VMs (Linux & Windows; VMware), cloud platforms (AWS, GCP, Azure), and containerized environments (Kubernetes)
- Diagnose and resolve issues across networking, Kubernetes, CDN, and traffic management layers (Akamai, waiting rooms, etc.)
Required Technical Skills & Experience :
Core Engineering & Operations :
- Strong hands-on experience in incident management and triage in production environments
- Proven ability to troubleshoot complex distributed systems under pressure
- Solid understanding of Linux systems administration (performance tuning, networking, NTP, etc.)
- Prior experience as an Incident Commander or in similar leadership roles during outages
Cloud & Infrastructure :
- Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2)
- Familiarity with GCP and/or Azure environments
- Proven experience operating in multi-cloud and hybrid environments
Containers & Orchestration :
- Experience troubleshooting Kubernetes clusters (pods, ingress, configuration issues)
- Understanding of containerized application architectures
DevOps & CI/CD :
- Strong knowledge of DevOps practices and CI/CD pipelines
- Hands-on experience with Harness, GitHub, and/or GitLab
Application & Technology Stack :
- Working knowledge of Java, Node.js, and React-based applications
- Understanding of database connectivity and dependencies across Oracle, MariaDB, and MSSQL
- Strong troubleshooting awareness (no DBA ownership required)
Networking :
- Strong foundational knowledge of TCP/IP, DNS, and HTTP(S)
- Experience with load balancing and network troubleshooting
- Ability to diagnose connectivity issues between services and databases
- 8 to 14 years of total IT experience.
- Minimum 7+ years of hands-on experience in Site Reliability Engineering (SRE) and/or DevOps.
- Proven experience in a Technical Lead/Team Lead role.
- Strong expertise in Linux administration and troubleshooting.
- Hands-on experience with automation and scripting (Shell, Python, PowerShell, etc.).
- Strong knowledge of DevOps tools and CI/CD pipelines.
- Experience working on at least one major cloud platform : Azure, AWS, or GCP.
- Experience managing production environments with a focus on availability, reliability, monitoring, and incident management.
- Strong troubleshooting, problem-solving, and root cause analysis skills.
- Notice Period Immediate only
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1649094