Posted on: 29/05/2026
Description :
Role & Responsibilities :
- Manage and maintain highly available, scalable, and reliable production infrastructure.
- Monitor system performance, uptime, latency, and application health across environments.
- Build and improve CI/CD pipelines for faster and safer deployments.
- Automate infrastructure provisioning, deployments, monitoring, and operational tasks.
- Troubleshoot production incidents, outages, and performance bottlenecks.
- Lead incident response, root cause analysis (RCA), and postmortem activities.
- Implement Infrastructure as Code (IaC) using tools like Terraform or Ansible.
- Manage Kubernetes clusters, container orchestration, and cloud-native platforms.
- Define and maintain SLI/SLO/SLA standards and reliability metrics.
- Optimize cloud infrastructure for scalability, performance, and cost efficiency.
- Collaborate with development, DevOps, security, and platform teams to improve system reliability.
- Ensure backup, disaster recovery, failover, and business continuity strategies are in place.
- Improve observability using monitoring, logging, and alerting tools.
- Participate in on-call rotations and production support activities.
- Mentor junior engineers and drive operational best practices.
Preferred Candidate Profile :
- Bachelors degree in Computer Science, Information Technology, or related field.
- 5 to 10 years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, or Platform Engineering.
- Strong hands-on experience with AWS, GCP, or Azure cloud platforms.
- Expertise in Kubernetes, Docker, and containerized environments.
- Strong experience with Terraform, Ansible, or other Infrastructure as Code tools.
- Proficiency in Linux administration and shell scripting.
- Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or ArgoCD.
- Knowledge of monitoring and observability tools like Prometheus, Grafana, Datadog, ELK, or Splunk.
- Good programming/scripting skills in Python, Go, or Bash.
- Strong understanding of networking concepts, security practices, and distributed systems.
- Experience handling production incidents and high-availability systems.
- Familiarity with microservices architecture and cloud-native technologies.
- Strong analytical, troubleshooting, and problem-solving abilities.
- Excellent communication and collaboration skills.
- Experience in 24x7 production support and on-call environments preferred.
Nice-to-Have Skills :
- Service Mesh (Istio/Linkerd)
- Kafka or distributed messaging systems
- FinOps / cloud cost optimization
- Chaos engineering
- MLOps or platform engineering exposure
- Multi-cloud infrastructure management
- DevSecOps practices
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1640124