HamburgerMenu
hirist

Site Reliability Engineer - Cloud Infrastructure

Astika Software Technologies
6 - 11 Years
Hyderabad

Posted on: 22/09/2026

Job Description

Key Responsibilities :

- Lead and drive Site Reliability Engineering (SRE) practices across production environments.

- Design, implement, and maintain highly available, scalable, and resilient infrastructure and services.

- Manage and optimize Kubernetes clusters and containerized workloads.

- Build and maintain cloud infrastructure across AWS and/or Azure.

- Develop and maintain Infrastructure as Code using Terraform.

- Design, maintain, and improve CI/CD pipelines for reliable and efficient software delivery.

- Implement automation to reduce manual operational effort and improve system reliability.

- Establish and monitor SLIs, SLOs, and error budgets for critical services.

- Develop and maintain monitoring, alerting, dashboards, and observability solutions using Prometheus and Grafana.

- Lead production incident management, including incident response, coordination, and resolution.

- Conduct detailed Root Cause Analysis (RCA) for production incidents and drive corrective and preventive actions.

- Identify system bottlenecks, reliability risks, and opportunities for performance and availability improvements.

- Implement proactive monitoring and capacity planning to prevent production issues.

- Establish and improve operational processes, runbooks, and reliability standards.

- Collaborate closely with Engineering, Development, QA, Security, and Product teams to improve application and infrastructure reliability.

- Mentor SRE/DevOps engineers and provide technical leadership on reliability and automation initiatives.

- Participate in on-call and production support activities as required.

Required Technical Skills :

- 8+ years of experience in SRE, DevOps, Infrastructure Engineering, or a related role.

- Strong hands-on experience with Kubernetes and Docker/containerized environments.

- Strong experience with AWS and/or Azure cloud platforms.

- Strong Linux administration and troubleshooting skills.

- Hands-on experience with Terraform and Infrastructure as Code.

- Strong understanding and experience with CI/CD practices and tools.

- Experience with Prometheus, Grafana, and modern monitoring/observability practices.

- Strong understanding of SLI, SLO, SLA, error budgets, and reliability engineering principles.

- Experience with production incident management, troubleshooting, and RCA.

- Strong scripting/automation skills using technologies such as Python, Bash, or similar.

- Good understanding of networking, DNS, load balancing, security, and distributed systems.

- Experience designing and operating highly available and fault-tolerant systems.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...