Posted on: 24/09/2026
Total Experience : 5 - 7 years
Key Responsibilities :
- Design, build, configure, and maintain scalable and reliable cloud infrastructure across AWS, Azure, and/or GCP.
- Deploy, manage, and troubleshoot Kubernetes-based environments.
- Perform software installation, configuration, upgrades, patching, and lifecycle management.
- Provide operational support, including incident management, problem management, troubleshooting, and root cause analysis.
- Build and support complex cloud infrastructure using Infrastructure as Code (IaC) and automation.
- Automate infrastructure provisioning, configuration, and operational tasks.
- Administer and troubleshoot Linux operating systems, including system performance, networking, processes, storage, and security.
- Develop automation scripts using Python and/or Linux Shell scripting.
- Develop and maintain Infrastructure as Code using Terraform, Ansible, or similar tools.
- Implement and maintain observability solutions covering metrics, logs, alerts, monitoring, and system health.
- Identify and resolve infrastructure, application, and platform reliability issues.
- Participate in on-call and production support activities as required.
- Collaborate with engineering, development, security, and operations teams to improve system reliability and operational efficiency.
- Apply Agile and DevOps principles across infrastructure and operations processes.
- Continuously improve automation, monitoring, deployment processes, and infrastructure reliability.
Required Skills & Experience:
- 5+ years of professional experience in Cloud Infrastructure, SRE, DevOps, or a related role.
- Strong hands-on experience with AWS and/or Azure and/or GCP.
- Advanced knowledge of Kubernetes and containerized environments.
- Strong experience with Linux administration and OS internals.
- Hands-on experience with Terraform, Ansible, or other Infrastructure as Code tools.
- Strong scripting/programming skills in Python and/or Linux Shell scripting.
- Experience with software installation, configuration, patching, and upgrades.
- Strong understanding of incident management, problem management, and production operations.
- Experience building and supporting cloud infrastructure programmatically.
- Knowledge of observability concepts, including metrics, logs, monitoring, and alerting.
- Understanding of DevOps, Agile, CI/CD, and automation practices.
- Strong troubleshooting and problem-solving skills.
- Ability to work effectively in a fast-paced, collaborative environment.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1674405