Posted on: 17/06/2026
Notice Period :
Immediate Joiners Only
Job Summary :
We are looking for an experienced Application Support & Reliability Engineer with a strong background in application support, operations, and reliability engineering. The ideal candidate will be responsible for ensuring application stability, troubleshooting production issues, supporting cloud and on-premises environments, and driving operational excellence. This role requires expertise in monitoring, automation, incident management, and observability, along with the ability to effectively manage critical incidents and collaborate with cross-functional teams.
Key Responsibilities :
Application Support & Operations :
- Provide application support and operational expertise across on-premises systems and cloud environments.
- Troubleshoot and resolve application, infrastructure, and system-related issues to ensure high availability and reliability.
- Investigate and diagnose production incidents and implement corrective actions.
- Support system performance and operational stability through proactive monitoring and issue resolution.
Monitoring & Observability :
- Utilize monitoring tools such as Nagios, Prometheus, and CloudWatch for system monitoring and alert management.
- Monitor application health and respond to alerts in a timely manner.
- Apply observability concepts including metrics, logs, and traces to identify and resolve issues effectively.
- Support continuous improvement of monitoring and operational processes.
Automation & Scripting :
- Develop and maintain automation scripts using Python, Shell, or similar technologies.
- Automate repetitive operational tasks to improve efficiency and reliability.
- Support operational processes through scripting and workflow automation.
Incident Management & Reliability :
- Manage incidents and support activities in alignment with incident management frameworks and ITIL processes.
- Participate in critical incident handling and drive timely resolution.
- Work closely with stakeholders to minimize downtime and ensure service continuity.
- Contribute to operational reliability initiatives and support best practices.
Cloud & DevOps Support :
- Support environments running on cloud platforms, with a preference for AWS.
- Understand cloud migration considerations and operational requirements.
- Work with CI/CD pipelines and DevOps tools to support deployments and release processes.
- Apply knowledge of observability and operational practices to improve system reliability.
Site Reliability Engineering Exposure :
- Contribute to reliability-focused initiatives and operational excellence efforts.
- Leverage exposure to SRE principles and practices to enhance system stability and support processes.
Required Skills & Qualifications :
Experience :
- Minimum 10+ years of experience in application support, operations, or reliability engineering.
Technical Skills :
- Strong troubleshooting skills across on-prem systems and cloud environments.
- Familiarity with Java, MySQL, and Python for debugging and support.
- Experience with monitoring tools such as Nagios, Prometheus, and CloudWatch, along with alert management.
- Experience with automation scripting using Python, Shell, or similar technologies.
- Knowledge of incident management frameworks and ITIL processes.
- Understanding of cloud platforms, with AWS preferred, and cloud migration considerations.
- Exposure to SRE principles and practices.
- Experience with CI/CD pipelines and DevOps tools.
- Knowledge of observability concepts including metrics, logs, and traces.
Soft Skills :
- Strong communication and collaboration skills.
- Ability to work under pressure and effectively manage critical incidents.
- Analytical mindset with a focus on problem-solving and operational excellence.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1645971