Posted on: 22/09/2026
Key Responsibilities :
- Lead and drive Site Reliability Engineering (SRE) practices across production environments.
- Design, implement, and maintain highly available, scalable, and resilient infrastructure and services.
- Manage and optimize Kubernetes clusters and containerized workloads.
- Build and maintain cloud infrastructure across AWS and/or Azure.
- Develop and maintain Infrastructure as Code using Terraform.
- Design, maintain, and improve CI/CD pipelines for reliable and efficient software delivery.
- Implement automation to reduce manual operational effort and improve system reliability.
- Establish and monitor SLIs, SLOs, and error budgets for critical services.
- Develop and maintain monitoring, alerting, dashboards, and observability solutions using Prometheus and Grafana.
- Lead production incident management, including incident response, coordination, and resolution.
- Conduct detailed Root Cause Analysis (RCA) for production incidents and drive corrective and preventive actions.
- Identify system bottlenecks, reliability risks, and opportunities for performance and availability improvements.
- Implement proactive monitoring and capacity planning to prevent production issues.
- Establish and improve operational processes, runbooks, and reliability standards.
- Collaborate closely with Engineering, Development, QA, Security, and Product teams to improve application and infrastructure reliability.
- Mentor SRE/DevOps engineers and provide technical leadership on reliability and automation initiatives.
- Participate in on-call and production support activities as required.
Required Technical Skills :
- 8+ years of experience in SRE, DevOps, Infrastructure Engineering, or a related role.
- Strong hands-on experience with Kubernetes and Docker/containerized environments.
- Strong experience with AWS and/or Azure cloud platforms.
- Strong Linux administration and troubleshooting skills.
- Hands-on experience with Terraform and Infrastructure as Code.
- Strong understanding and experience with CI/CD practices and tools.
- Experience with Prometheus, Grafana, and modern monitoring/observability practices.
- Strong understanding of SLI, SLO, SLA, error budgets, and reliability engineering principles.
- Experience with production incident management, troubleshooting, and RCA.
- Strong scripting/automation skills using technologies such as Python, Bash, or similar.
- Good understanding of networking, DNS, load balancing, security, and distributed systems.
- Experience designing and operating highly available and fault-tolerant systems.
Did you find something suspicious?
Posted by
ASTIKA SOFTWARE TECHNOLOGIES PRIVATE LIMITED
Decision Maker at Astika Software Technologies
Last Active: 22 Sep 2026
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1673574