HamburgerMenu
hirist

Lead Site Reliability Engineer

HR Works Consultancy
8 - 14 Years
Multiple Locations

Posted on: 27/06/2026

Job Description

Job Description :

Incident Response & Leadership :

- Lead and mentor a team of 5-6 SREs in daily operations and incident response activities

- Act as the first responder to production alerts, rapidly assessing severity and initiating mitigation

- Serve as Incident Commander during major incidents, leading bridge calls with clarity and decisiveness

- Drive root cause isolation within 30 minutes for critical incidents whenever possible

- Communicate effectively with engineering, product, and leadership teams during high-pressure situations

- Maintain strong presence and ownership on incident bridges with confident decision-making

Team Management & Operations :

- Oversee daily activities and coordinate closely with client leads and managers

- Prepare weekly and monthly KPI reports on team performance and reliability metrics

- Drive continuous improvement initiatives across the team

Proactive Reliability Engineering :

- Identify patterns, trends, and signals to prevent incidents before they occur

- Continuously improve alert quality, reduce noise, and increase signal fidelity

- Partner with engineering teams to enhance system resilience and reliability

Automation & Toil Reduction :

- Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows

- Build and improve tooling across incident response, observability, and operations

- Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value

Platform & Systems Support :

- Troubleshoot across hybrid ecosystems including on-prem VMs (Linux & Windows; VMware), cloud platforms (AWS, GCP, Azure), and containerized environments (Kubernetes)

- Diagnose and resolve issues across networking, Kubernetes, CDN, and traffic management layers (Akamai, waiting rooms, etc.)

Required Technical Skills & Experience :

Core Engineering & Operations :

- Strong hands-on experience in incident management and triage in production environments

- Proven ability to troubleshoot complex distributed systems under pressure

- Solid understanding of Linux systems administration (performance tuning, networking, NTP, etc.)

- Prior experience as an Incident Commander or in similar leadership roles during outages

Cloud & Infrastructure :

- Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2)

- Familiarity with GCP and/or Azure environments

- Proven experience operating in multi-cloud and hybrid environments

Containers & Orchestration :

- Experience troubleshooting Kubernetes clusters (pods, ingress, configuration issues)

- Understanding of containerized application architectures

DevOps & CI/CD :

- Strong knowledge of DevOps practices and CI/CD pipelines

- Hands-on experience with Harness, GitHub, and/or GitLab

Application & Technology Stack :

- Working knowledge of Java, Node.js, and React-based applications

- Understanding of database connectivity and dependencies across Oracle, MariaDB, and MSSQL

- Strong troubleshooting awareness (no DBA ownership required)

Networking :

- Strong foundational knowledge of TCP/IP, DNS, and HTTP(S)

- Experience with load balancing and network troubleshooting

- Ability to diagnose connectivity issues between services and databases

- 8 to 14 years of total IT experience.

- Minimum 7+ years of hands-on experience in Site Reliability Engineering (SRE) and/or DevOps.

- Proven experience in a Technical Lead/Team Lead role.

- Strong expertise in Linux administration and troubleshooting.

- Hands-on experience with automation and scripting (Shell, Python, PowerShell, etc.).

- Strong knowledge of DevOps tools and CI/CD pipelines.

- Experience working on at least one major cloud platform : Azure, AWS, or GCP.

- Experience managing production environments with a focus on availability, reliability, monitoring, and incident management.

- Strong troubleshooting, problem-solving, and root cause analysis skills.

- Notice Period Immediate only

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...