Posted on: 10/04/2026
Description :
Key Responsibilities :
- Design and maintain monitoring, logging, and alerting systems across environments
- Lead incident response, RCA, and post-mortem analysis
- Implement and test disaster recovery strategies
- Collaborate with teams to define and maintain SLAs
- Optimize cloud infrastructure for performance, reliability, and cost
- Build automation for deployment, scaling, and recovery
- Manage infrastructure using Terraform, GitLab CI/CD, and Kubernetes
- Participate in on-call rotations
Required Skills & Experience :
- Strong experience in SRE/DevOps roles (5+ years)
- Proficiency in AWS (EC2, EKS, RDS, CloudWatch, Cognito)
- Hands-on experience with Kubernetes in production environments
- Expertise in Infrastructure as Code (Terraform/CloudFormation)
- Strong scripting skills in Python, Bash, or Shell
- Experience with Chef and Ansible
- Strong understanding of monitoring tools (Prometheus, Grafana, ELK)
- Experience with PostgreSQL or other RDBMS and replication
- Knowledge of networking, load balancing, and security best practices
- Experience with CI/CD pipelines and GitOps workflows
Key Skills :
- Linux Administration , Docker , Kubernetes , Nginx , Redis , Kafka
- CICD , Jenkins , Git , Ansible , Shell Scripting
- ELK , Grafana , Prometheus , Networking , Firewalls , WAF , Akamai
- RCA , Capacity Planning , RDBMS , Java
If you are passionate about scalability, automation, and reliability, we would love to connect.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1627513