Posted on: 14/07/2026
Job Description :
We are looking for an experienced Site Reliability Engineer (Azure) to support a cloud-native, Azure-based data platform focused on delivering business insights and KPI dashboards. The ideal candidate should have strong expertise in Azure cloud services, observability, monitoring, and production support, with the ability to ensure the reliability and performance of distributed applications and data pipelines.
Key Responsibilities :
- Ensure the reliability, availability, and performance of Azure-based applications and data platforms.
- Monitor and support end-to-end data processing pipelines, APIs, and cloud services to ensure timely delivery of business insights.
- Implement and optimize observability, monitoring, logging, and alerting across production environments.
- Investigate production incidents, perform root cause analysis, and drive issue resolution through to closure.
- Collaborate with engineering, DevOps, and data teams to improve platform reliability, scalability, and operational efficiency.
- Develop and maintain monitoring dashboards, health checks, and automated alerting mechanisms.
- Troubleshoot Azure services, REST APIs, and event-driven workflows to minimize downtime and improve system resilience.
- Perform minor development and debugging tasks using C# to support production issues and system enhancements.
- Support hypercare activities during releases and major deployments.
- Contribute to automation initiatives using Infrastructure as Code (IaC) and operational best practices.
Required Skills :
- 5-10 years of experience in Site Reliability Engineering (SRE), Production Support, or Cloud Operations.
- Strong expertise in Microsoft Azure cloud services and cloud-native architectures.
- Hands-on experience with observability and monitoring tools such as Azure Application Insights, Elastic Stack (ELK), or similar platforms.
- Good understanding of REST APIs and event-driven architectures, including Azure Service Bus.
- Proficiency in C# for troubleshooting, debugging, and minor application enhancements.
- Strong experience with incident management, root cause analysis, and production support.
- Excellent analytical, communication, and stakeholder management skills.
- Strong ownership mindset with the ability to manage production issues end-to-end.
Preferred Skills :
- Experience with Terraform and Infrastructure as Code (IaC).
- Familiarity with Azure Databricks, AI/ML pipelines, and data engineering workflows.
- Knowledge of React.js for front-end troubleshooting.
- Experience with distributed systems and event-driven application architectures.
- Exposure to CI/CD pipelines and DevOps practices.
Education (Mandatory) :
- Full-time B.E./B.Tech, MCA, M.Tech, MS, or M.Sc in Computer Science, Information Technology, Engineering, or a related discipline.
Did you find something suspicious?
Posted by
Menaka Sulladmath
Managing Director at BUSINESS FUNDAMENTAL CONSULTING INDIA PRIVATE
Last Active: 11 Aug 2026
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1654032