HamburgerMenu
hirist

Senior Reliability Engineer - Cloud Infrastructure

Good Co
5 - 10 Years
Anywhere in India/Multiple Locations

Posted on: 03/08/2026

Job Description

Role & Responsibilities :

- Own the reliability, availability, scalability, and performance of production systems and critical services.

- Design, implement, and maintain highly available cloud infrastructure across AWS/Azure/GCP environments.

- Define and improve SLOs, SLIs, and SLAs to measure and enhance system reliability.

- Lead incident response, troubleshoot complex production issues, perform root cause analysis (RCA), and drive preventive actions.

- Build and maintain automation for infrastructure provisioning, deployments, monitoring, and operational workflows.

- Manage and optimize Kubernetes clusters, containerized workloads, and cloud-native architectures.

- Develop and maintain CI/CD pipelines to enable safe, frequent, and reliable software releases.

- Implement observability solutions using tools such as Prometheus, Grafana, Datadog, New Relic, Splunk, or ELK.

- Improve system performance through capacity planning, load testing, scalability improvements, and tuning.

- Establish disaster recovery strategies, backup processes, and business continuity plans.

- Automate repetitive operational tasks using scripting and infrastructure-as-code practices.

- Collaborate with engineering, security, and product teams to improve system resilience.

- Participate in on-call rotations and provide technical leadership during high-priority incidents.

Preferred Candidate Profile :

- Bachelors degree in Computer Science, Information Technology, Engineering, or a related field.

- 5-10 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Production Engineering roles.

- Hands-on expertise with at least one major cloud platform:

1. AWS

2. Microsoft Azure

3. Google Cloud Platform (GCP)

- Strong knowledge of:

1. Linux administration

2. Networking fundamentals (TCP/IP, DNS, HTTP, load balancing)

3. Distributed systems concepts

4. System troubleshooting and performance optimization

- Proficiency in scripting/programming languages:

1. Python

2. Bash

3. Go (preferred)

- Experience with monitoring and observability platforms:

1. Prometheus

2. Grafana

3. Datadog

4. Splunk

5. ELK Stack

Strong understanding of :

1. Incident management

2. Root cause analysis

3. Error budgets

4. Capacity planning

5. Reliability metrics

- Familiarity with security best practices, IAM, vulnerability management, and compliance requirements.

- Excellent communication skills with the ability to work with globally distributed teams.

- Ability to independently manage production responsibilities in a remote environment.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...