Posted on: 26/06/2026
Job Description :
We are seeking a highly skilled SRE AWS DevOps Engineer to build, automate, and manage scalable cloud infrastructure while ensuring high availability, reliability, and performance of business-critical applications. The ideal candidate will have strong expertise in AWS, DevOps practices, infrastructure automation, container orchestration, monitoring, and incident management.
Key Responsibilities :
- Design, implement, and manage secure, scalable, and highly available cloud infrastructure on AWS.
- Build and maintain Infrastructure as Code (IaC) solutions using tools such as Terraform, CloudFormation, or Ansible.
- Develop, optimize, and support CI/CD pipelines to enable efficient and reliable application deployments.
- Manage containerized environments using Kubernetes and related cloud-native technologies.
- Implement Site Reliability Engineering (SRE) practices to improve system availability, resiliency, scalability, and operational efficiency.
- Monitor application and infrastructure performance using observability tools and proactively identify potential issues.
- Automate operational processes, deployments, monitoring, incident response, and remediation activities.
- Collaborate with development, security, and infrastructure teams to enhance platform reliability and performance.
- Troubleshoot and resolve issues related to infrastructure, deployments, networking, and application performance.
- Conduct root cause analysis (RCA) for production incidents and implement preventive measures.
- Ensure adherence to security, compliance, change management, and operational governance standards.
- Manage backup, disaster recovery, capacity planning, and business continuity initiatives.
- Maintain operational documentation, runbooks, and standard operating procedures.
Required Skills & Experience :
- 5-10 years of experience in Site Reliability Engineering, DevOps Engineering, Cloud Engineering, or Platform Engineering.
- Strong hands-on experience with AWS services including EC2, VPC, IAM, S3, RDS, EKS, CloudWatch, Route 53, and related services.
- Expertise in Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation.
- Strong experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
- Hands-on experience with Kubernetes and containerization technologies such as Docker.
- Experience with monitoring, logging, and observability tools including Prometheus, Grafana, ELK, Datadog, or CloudWatch.
- Strong understanding of Linux systems administration, networking, and cloud security best practices.
- Experience with incident management, problem management, and production support processes.
- Proficiency in scripting and automation using Python, Shell, or Bash.
- Strong troubleshooting, analytical, and problem-solving skills.
Preferred Skills :
- Experience with multi-cloud environments (AWS, Azure, GCP).
- Knowledge of SRE concepts such as SLIs, SLOs, error budgets, and reliability engineering practices.
- Familiarity with service mesh technologies and cloud-native architectures.
- Experience implementing DevSecOps practices and security automation.
- AWS, Kubernetes, Terraform, or DevOps-related certifications are an advantage.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1648737