HamburgerMenu
hirist

Lead Site Reliability Engineer

BP Infosystem LLP
9 - 14 Years
Multiple Locations

Posted on: 16/06/2026

Job Description

Role Overview :

- As a Lead Site Reliability Engineer, you will serve as the technical anchor for our mission-critical infrastructure, ensuring high availability, scalability, and performance across our global cloud footprint.

- You will lead a high-performing team of engineers to bridge the gap between development and operations, fostering a culture of automation and proactive system health.

- By collaborating closely with product engineering teams and senior stakeholders, you will architect robust solutions that minimize downtime and optimize resource utilization.

- Your work directly impacts the end-user experience by ensuring seamless service delivery and maintaining the integrity of our production environments, ultimately driving business reliability and operational excellence.

Key Responsibilities :

Incident Response & Leadership :

- Lead and mentor a team of 5-6 SREs in daily operations and incident response activities.

- Act as the first responder to production alerts, rapidly assessing severity and initiating mitigation.

- Serve as Incident Commander during major incidents, leading bridge calls with clarity and decisiveness.

- Drive root cause isolation within 30 minutes for critical incidents whenever possible.

- Communicate effectively with engineering, product, and leadership teams during high-pressure situations.

- Maintain strong presence and ownership on incident bridges with confident decision-making.

Team Management & Operations :

- Oversee daily activities and coordinate closely with client leads and managers.

- Prepare weekly and monthly KPI reports on team performance and reliability metrics.

- Drive continuous improvement initiatives across the team.

Proactive Reliability Engineering :

- Identify patterns, trends, and signals to prevent incidents before they occur.

- Continuously improve alert quality, reduce noise, and increase signal fidelity.

- Partner with engineering teams to enhance system resilience and reliability.

Automation & Toil Reduction :

- Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows.

- Build and improve tooling across incident response, observability, and operations.

- Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value.

Platform & Systems Support :

- Troubleshoot across hybrid ecosystems including on-prem VMs (Linux & Windows; VMware), cloud platforms (AWS, GCP, Azure), and containerized environments (Kubernetes).

- Diagnose and resolve issues across networking, Kubernetes, CDN, and traffic management layers (Akamai, waiting rooms, etc.).

Required Skillset :

Core Engineering & Operations :

- Strong hands-on experience in incident management and triage in production environments.

- Proven ability to troubleshoot complex distributed systems under pressure.

- Solid understanding of Linux systems administration (performance tuning, networking, NTP, etc.).

- Prior experience as an Incident Commander or in similar leadership roles during outages.

Cloud & Infrastructure :

- Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2).

- Familiarity with GCP and/or Azure environments.

- Proven experience operating in multi-cloud and hybrid environments.

- Experience troubleshooting Kubernetes clusters (pods, ingress, configuration issues).

- Understanding of containerized application architectures.

DevOps & CI/CD :

- Strong knowledge of DevOps practices and CI/CD pipelines.

- Hands-on experience with Harness, GitHub, and/or GitLab.

Application & Technology Stack :

- Understanding of database connectivity and dependencies across Oracle, MariaDB, and MSSQL.

Networking :

- Strong foundational knowledge of TCP/IP, DNS, and HTTP(S).

- Experience with load balancing and network troubleshooting.

- Ability to diagnose connectivity issues between services and databases.

Preferred Qualifications :

- Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications.

- Familiarity with Akamai CDN and traffic management tools.

- Experience in high-volume, high-availability production environments.

- Track record of leading teams in fast-paced, incident-driven environments.

Working Hours :

- Shift 1 : 3 :00 AM - 12 :30 PM IST (5 :30 PM - 3 :00 AM EST)

- Shift 2 : 11 :00 AM - 8 :30 PM IST (1 :30 AM - 11 :00 AM EST)

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...