Posted on: 30/06/2026
Role : Site Reliability Engineer (SRE)
Experience : 8-12 Years
Location : Mumbai
Work Mode : Onsite/Hybrid (as per business requirements)
About the Role :
We are looking for an experienced Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and resilient infrastructure.
The ideal candidate should have strong expertise in DevOps practices, Infrastructure as Code (IaC), container orchestration, and cloud-native technologies to ensure platform reliability, performance, and operational excellence.
Key Responsibilities :
- Design, implement, and maintain highly available and scalable production infrastructure.
- Automate infrastructure provisioning and management using Terraform.
- Manage containerized applications using Docker and Kubernetes.
- Improve system reliability, performance, availability, and scalability through automation and monitoring.
- Develop and maintain CI/CD pipelines to enable efficient software delivery.
- Monitor infrastructure health, troubleshoot production issues, and perform root cause analysis.
- Collaborate with Development, DevOps, Security, and Infrastructure teams to enhance platform stability.
- Implement observability solutions including logging, monitoring, alerting, and incident management.
- Optimize infrastructure costs, system performance, and deployment processes.
- Participate in production support, on-call rotations, and incident response activities.
- Ensure infrastructure follows security, compliance, and operational best practices.
Required Skills :
- 812 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
- Strong hands-on experience in Site Reliability Engineering (SRE) principles and practices.
- Expertise in Terraform for Infrastructure as Code (IaC).
- Strong knowledge of Docker and Kubernetes.
- Solid understanding of DevOps methodologies and automation.
- Experience with CI/CD tools and deployment automation.
- Strong understanding of Linux systems, networking, and troubleshooting.
- Experience with monitoring, logging, and observability tools.
- Knowledge of cloud platforms such as AWS, Azure, or GCP is an advantage.
- Strong scripting skills using Bash, Python, or Go are preferred.
- Excellent analytical, troubleshooting, and problem-solving skills.
Good to Have :
- Experience with GitOps practices and tools.
- Knowledge of service mesh technologies such as Istio or Linkerd.
- Experience with Prometheus, Grafana, ELK/OpenSearch, or similar monitoring tools.
- Exposure to cloud security and container security best practices.
- Experience with microservices architecture and distributed systems.
Preferred Candidate Profile :
- Strong ownership mindset with a focus on reliability and automation.
- Ability to troubleshoot complex production issues under pressure.
- Excellent communication and collaboration skills.
- Passion for automation, platform engineering, and continuous improvement.
- Experience working in Agile and DevOps environments.
Did you find something suspicious?
Posted by
HR
HR Associate at Arting Digital
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1650136