Posted on: 16/09/2026
Key Responsibilities :
1. Cloud Operations & Incident Management :
- Provide hands-on support, maintenance, and troubleshooting for high-availability production environments across AWS and GCP.
- Lead incident management workflows, performing root-cause analysis (RCA), resolving critical production outages, and driving post-mortem reviews.
- Manage production uptime, capacity planning, performance tuning, and service availability SLAs.
2. Configuration Management & Infrastructure as Code (IaC) :
- Utilize SaltStack (Mandatory) to enforce configuration standards, automate state management, and orchestrate large-scale node environments.
- Provision, update, and manage cloud infrastructure using Terraform following IaC best practices.
- Maintain consistency and security postures across multi-cloud infrastructure environments.
3. Observability, Automation & Scripting :
- Build, maintain, and enhance enterprise monitoring, logging, and observability platforms (e.g., Datadog, Prometheus, Grafana, CloudWatch).
- Automate operational tasks and routine maintenance using scripting languages (Python, Bash, or Shell).
- Implement automated proactive alerting mechanisms to reduce MTTR (Mean Time to Resolution).
4. Technical Governance & Collaboration :
- Author comprehensive system documentation, standard operating procedures (SOPs), and architectural runbooks.
- Collaborate closely with DevOps and Engineering teams to transition new applications into production smoothly.
Technical Qualifications :
Required Experience & Skills :
Experience : 6+ years in Cloud Operations, Cloud Infrastructure, or Site Reliability Engineering (SRE).
Multi-Cloud Support : Proven hands-on production support across both AWS and GCP.
Configuration Management : Expert-level, mandatory proficiency in SaltStack for automated server configuration and orchestration.
Infrastructure as Code : Hands-on proficiency implementing IaC using Terraform.
Operational Readiness : Extensive experience in production support, high-severity incident management, and complex system troubleshooting.
Observability : Practical experience deploying and managing modern monitoring, logging, and metrics platforms.
Automation : Strong automation mindset with proven scripting skills (Python, Bash, etc.).
Soft Skills : Excellent written/verbal communication and documentation capabilities.
Preferred / Nice-to-Have Skills :
- Infrastructure provisioning with AWS CloudFormation.
- Hands-on Kubernetes Administration and container orchestration.
- Experience with DevOps & CI/CD pipeline development (GitLab CI, GitHub Actions, Jenkins).
- Strong Linux Systems Administration foundation (RHEL/CentOS/Ubuntu).
- Database Lifecycle Management : Experience handling database upgrades, maintenance, and failovers (Relational & NoSQL).
- Proven involvement in Cloud Migration initiatives.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1671663