HamburgerMenu
hirist

Job Description

Role Overview :

As a seasoned Site Reliability Engineer, you will be instrumental in ensuring the unwavering availability, performance, and scalability of our critical production systems. This role demands a proactive approach to operational excellence, working closely with our development, product, and infrastructure teams to embed reliability principles throughout the software development lifecycle.


You will champion a culture of automation and continuous improvement, directly impacting the stability of our services, enhancing user trust, and safeguarding business continuity for our diverse customer base.

Key Responsibilities :

- Lead the architectural design and implementation of highly resilient and scalable infrastructure solutions, focusing on Service Reliability and performance optimization for critical applications.

- Drive the development and enforcement of robust Disaster Recovery and IT Disaster Recovery plans, conducting regular drills and ensuring swift, effective recovery strategies to minimize downtime.

- Define, monitor, and optimize Service Level Indicators (SLIs) and Service Level Objectives (SLOs), establishing `Error Budgets` to balance innovation with operational stability and meet stringent `SLA` commitments.

- Spearhead Incident Management processes, leading post-incident reviews to identify root causes, implement corrective actions, and prevent recurrence, fostering a learning culture.

- Develop and implement advanced Troubleshooting methodologies and tools to quickly diagnose and resolve complex system issues across distributed environments.

- Automate operational tasks, infrastructure provisioning, and deployment pipelines using modern SRE practices, reducing manual toil and improving system efficiency.

- Provide expert guidance and mentorship in System Administration best practices, cloud infrastructure, and observability tools to cross-functional teams, elevating overall technical capabilities.

- Collaborate with engineering teams to design and build systems that are inherently observable, scalable, and maintainable from inception.

- Evaluate and integrate new technologies and tools to enhance our reliability posture, security, and operational efficiency.

Required Skillset :

- Demonstrated expertise in designing, building, and operating large-scale, highly available distributed systems, with a deep understanding of Site Reliability principles.

- Proven ability to lead and execute complex Disaster Recovery and IT Disaster Recovery initiatives, including strategy formulation, testing, and execution.

- Strong analytical skills to define, track, and interpret Service Level Indicators (SLIs), manage Error Budgets, and ensure SLA adherence.

- Exceptional Troubleshooting and problem-solving capabilities for intricate production issues across diverse technology stacks and cloud environments.

- Extensive experience in System Administration for Linux-based systems, including networking, security, and performance tuning.

- Proficient in scripting and automation (e.g., Python, Go, Shell) and experience with infrastructure-as-code tools (e.g., Terraform, Ansible).

- Strong leadership and Incident Management skills, with the ability to remain calm and decisive under pressure during critical outages.

- Excellent communication and interpersonal skills, capable of articulating complex technical concepts to both technical and non-technical stakeholders.

- Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field from a reputable institution.

- Adaptability to work in a dynamic, fast-paced environment based out of our Navi Mumbai office, collaborating effectively with global teams.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...