Posted on: 16/09/2026
Job Description :
About the Role :
We are looking for a Site Reliability Engineer / Platform SRE who enjoys working with cloud infrastructure, Kubernetes and production environments.
The ideal candidate should be comfortable troubleshooting real-world production issues, improving platform reliability, monitoring applications and working closely with Platform and DevOps teams.
This is not a coding-heavy role. We are looking for someone who can understand existing code and scripts, troubleshoot issues and make basic changes when required. AI-assisted tools can be used for more complex coding requirements.
Key Responsibilities :
- Manage and support Azure-based cloud infrastructure and services.
- Work with Kubernetes for deployments, troubleshooting, scaling and production support.
- Monitor production environments and investigate application/infrastructure issues.
- Analyze logs, metrics and alerts to identify issues and perform root-cause analysis (RCA).
- Work closely with Platform Engineering and DevOps teams to maintain reliable production environments.
- Support and improve CI/CD pipelines and deployment processes.
- Work with Apache Flink and data-processing environments where required.
- Support ETL/ELT pipelines and data-platform workflows.
- Contribute to monitoring and observability using tools such as Datadog, Dynatrace or similar platforms.
- Participate in incident management and help improve system reliability.
- Support vulnerability identification, remediation and other basic security-related activities.
- Assist with API performance/load testing and identifying performance bottlenecks.
- Use automation and AI-assisted tools to improve operational efficiency.
Required Skills :
- Hands-on experience with Microsoft Azure.
- Strong practical experience with Kubernetes.
- Good understanding of SRE / Platform Engineering concepts.
- Strong production troubleshooting and RCA experience.
- Experience with CI/CD and deployment processes.
- Good understanding of monitoring, logging and observability.
- Exposure to Datadog, Dynatrace or similar observability tools.
- Experience with Apache Flink is preferred.
- Understanding of ETL/ELT and data pipelines.
- Experience working with Platform/DevOps team.
Role Details :
- Role : Site Reliability Engineer
- Industry Type : Miscellaneous
- Department : Engineering - Software & QA
- Employment Type : Full Time, Permanent
- Role Category : DevOps
Education :
- UG : B.Tech / B.E. in Any Specialization, Any Graduate
Key Skills :
- Skills highlighted with are preferred keyskills : Site Reliability Engineering, Kubernetes, Datadog.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1671922