Posted on: 27/07/2026
Role Overview :
We are seeking a seasoned Site Reliability Engineer to join our high-performance engineering team in Pune. In this role, you will act as the bridge between development and operations, ensuring our mission-critical platforms remain resilient, scalable, and highly available.
You will collaborate closely with cross-functional product teams, architects, and stakeholders to design robust infrastructure solutions that minimize downtime and optimize system performance. By championing reliability engineering principles, you will directly influence the end-user experience, ensuring that our digital services deliver seamless performance for our global customer base.
Key Responsibilities :
- Architect and maintain scalable CI/CD pipelines to accelerate deployment cycles while ensuring rigorous quality gates for production releases.
- Lead incident management and root cause analysis efforts, utilizing ITSM best practices to resolve complex system bottlenecks and prevent recurring outages.
- Automate manual operational tasks through advanced scripting to reduce toil and improve overall system efficiency.
- Implement comprehensive monitoring and logging strategies using Splunk to gain deep visibility into system health and performance metrics.
- Partner with engineering teams to define and track Service Level Objectives (SLOs) and Error Budgets, ensuring alignment between technical performance and business goals.
Required Skillset :
- Demonstrated expertise in managing complex Linux environments, with a deep understanding of kernel tuning, performance troubleshooting, and system security.
- Proven ability to develop complex automation scripts using Python and Shell scripting to streamline infrastructure-as-code and configuration management.
- Strong proficiency in CI/CD orchestration and DevOps methodologies, with a track record of optimizing deployment workflows in large-scale distributed systems.
- Advanced analytical skills in log aggregation and data visualization using Splunk to proactively identify and mitigate potential system failures.
- Exceptional communication skills, with the ability to articulate technical risks to non-technical stakeholders and lead cross-functional initiatives effectively.
- A minimum of 6 to 14 years of professional experience in SRE or DevOps roles, demonstrating a consistent track record of managing high-traffic production environments.
- Ability to thrive in a hybrid work environment in Pune, maintaining high levels of collaboration and accountability within a distributed team structure.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1657920