Posted on: 17/08/2026
About the Role :
We are looking for an experienced Cloud Architect / DevOps Architect with strong hands-on expertise in cloud infrastructure, DevOps, resiliency and observability.
The role requires someone who can architect and engineer highly available, resilient and observable cloud platforms for enterprise environments. The ideal candidate should have strong AWS architecture experience, combined with hands-on knowledge of DevOps automation, Kubernetes, Infrastructure as Code, monitoring and Site Reliability Engineering practices.
This is a technically hands-on role requiring the ability to design solutions as well as work closely with engineering teams during implementation, troubleshooting and production stabilization.
Key Responsibilities :
- Design and implement scalable, secure and highly available cloud architectures, primarily on AWS.
- Architect cloud environments across compute, networking, storage, security, monitoring and automation services.
- Design High Availability (HA), Disaster Recovery (DR), failover and replication strategies for mission-critical applications.
- Build resiliency into cloud platforms through fault-tolerant architecture, automated recovery and recovery orchestration.
- Implement and improve observability and monitoring across applications, infrastructure and Kubernetes environments.
- Work with tools such as Prometheus, Grafana, ELK/OpenSearch, Splunk, Datadog, Dynatrace, New Relic and OpenTelemetry.
- Design and enhance CI/CD and DevOps automation using Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps or similar tools.
- Provision and manage infrastructure using Terraform, CloudFormation, Ansible and other Infrastructure as Code/automation tools.
- Design and support containerized workloads using Docker and Kubernetes.
- Apply SRE and reliability engineering principles to improve system availability, performance and operational stability.
- Participate in incident management, troubleshooting, Root Cause Analysis (RCA) and production reliability improvements.
- Support chaos testing, failure simulations and recovery validation to identify weaknesses before they impact production.
- Ensure cloud environments follow enterprise standards around security, networking, compliance and governance.
Must-Have Skills :
- 8+ years of experience in Cloud Architecture, DevOps, SRE or Platform Engineering.
- Strong hands-on experience with AWS cloud architecture is preferred.
- Deep understanding of resiliency, reliability and High Availability architecture.
- Strong experience designing HA/DR, failover, replication and disaster recovery solutions.
- Hands-on experience with observability and monitoring platforms such as Prometheus, Grafana, ELK/OpenSearch, Splunk, Datadog, Dynatrace, New Relic or OpenTelemetry.
- Strong experience with DevOps and CI/CD automation.
- Experience with Terraform, CloudFormation, Ansible or similar infrastructure automation tools.
- Hands-on knowledge of Docker and Kubernetes.
- Good understanding of SRE practices, incident management and Root Cause Analysis.
- Strong understanding of cloud networking, security, governance and enterprise production environments.
What We Are Specifically Looking For :
- We are particularly interested in candidates whose recent experience demonstrates strong hands-on ownership of cloud resiliency and observability, rather than profiles focused primarily on cloud migration or DevOps pipeline management.
- The ideal candidate should be able to architect a production cloud environment for high availability and disaster recovery, define how it will be monitored and observed, identify potential failure scenarios and design automated mechanisms to recover from them.
- Strong AWS + DevOps + Resiliency + Observability experience will be particularly relevant for this opportunity.
Important :
- This is an urgent requirement. Candidates who are available immediately or can join within a maximum of 15 days will be considered.
The job is for:
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1663715