HamburgerMenu
hirist

EMBARK - Senior Site Reliability Engineer - Cloud/Kubernetes

Embark Business Solutions
3 - 7 Years
Bangalore

Posted on: 16/09/2026

Job Description

Job Description :

About the Role :

We are looking for a Site Reliability Engineer / Platform SRE who enjoys working with cloud infrastructure, Kubernetes and production environments.

The ideal candidate should be comfortable troubleshooting real-world production issues, improving platform reliability, monitoring applications and working closely with Platform and DevOps teams.

This is not a coding-heavy role. We are looking for someone who can understand existing code and scripts, troubleshoot issues and make basic changes when required. AI-assisted tools can be used for more complex coding requirements.

Key Responsibilities :

- Manage and support Azure-based cloud infrastructure and services.

- Work with Kubernetes for deployments, troubleshooting, scaling and production support.

- Monitor production environments and investigate application/infrastructure issues.

- Analyze logs, metrics and alerts to identify issues and perform root-cause analysis (RCA).

- Work closely with Platform Engineering and DevOps teams to maintain reliable production environments.

- Support and improve CI/CD pipelines and deployment processes.

- Work with Apache Flink and data-processing environments where required.

- Support ETL/ELT pipelines and data-platform workflows.

- Contribute to monitoring and observability using tools such as Datadog, Dynatrace or similar platforms.

- Participate in incident management and help improve system reliability.

- Support vulnerability identification, remediation and other basic security-related activities.

- Assist with API performance/load testing and identifying performance bottlenecks.

- Use automation and AI-assisted tools to improve operational efficiency.

Required Skills :

- Hands-on experience with Microsoft Azure.

- Strong practical experience with Kubernetes.

- Good understanding of SRE / Platform Engineering concepts.

- Strong production troubleshooting and RCA experience.

- Experience with CI/CD and deployment processes.

- Good understanding of monitoring, logging and observability.

- Exposure to Datadog, Dynatrace or similar observability tools.

- Experience with Apache Flink is preferred.

- Understanding of ETL/ELT and data pipelines.

- Experience working with Platform/DevOps team.

Role Details :

- Role : Site Reliability Engineer

- Industry Type : Miscellaneous

- Department : Engineering - Software & QA

- Employment Type : Full Time, Permanent

- Role Category : DevOps

Education :

- UG : B.Tech / B.E. in Any Specialization, Any Graduate

Key Skills :

- Skills highlighted with are preferred keyskills : Site Reliability Engineering, Kubernetes, Datadog.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...