HamburgerMenu
hirist

Site Reliability Engineer

Mpowerment Resources
6 - 8 Years
rupee20-25 LPA
Kolkata

Posted on: 09/09/2026

Job Description

Location : Kolkata

Work Mode : Hybrid / Work from Office

Experience : 6 - 8 Years

Joining : Immediate Joiners Preferred

About the Role :

We are looking for an experienced Site Reliability Engineer (SRE) to join our Infrastructure & Platform Engineering - AI Product Services team.

The ideal candidate will be responsible for ensuring the reliability, scalability, availability, security, and performance of production infrastructure supporting AI/ML workloads. You will work closely with engineering, DevOps, platform, and product teams to build and operate highly scalable and resilient systems.

This role is particularly suited for professionals with strong experience in cloud infrastructure, Kubernetes, automation, production operations, and AI/ML workloads.

Key Responsibilities :

- Design, implement, and maintain highly reliable, scalable, and resilient production infrastructure.

- Manage and optimize cloud infrastructure across AWS and/or Azure.

- Deploy, operate, and troubleshoot containerized workloads using Kubernetes and Docker.

- Build and maintain CI/CD pipelines and deployment automation.

- Monitor system health, performance, availability, and capacity across production environments.

- Implement effective monitoring, logging, tracing, alerting, and observability solutions.

- Support and optimize infrastructure for ML/LLM workloads, GPU infrastructure, and model serving.

- Participate in incident response, troubleshooting, and Root Cause Analysis (RCA) for production issues.

- Identify performance bottlenecks and implement solutions to improve system reliability and scalability.

- Implement and maintain infrastructure security practices, including secrets management and vulnerability management.

- Automate infrastructure provisioning and configuration using Terraform and/or CloudFormation.

- Establish and improve SRE best practices around availability, reliability, scalability, and operational excellence.

- Collaborate with development and platform engineering teams to improve system architecture and deployment processes.

Required Skills & Experience :

- 6 - 8 years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, or a related role.

- Strong hands-on experience with AWS and/or Azure.

- Strong knowledge of Kubernetes and Docker.

- Experience with CI/CD tools and deployment automation.

- Strong understanding of production reliability, scalability, and high availability.

- Experience with monitoring, logging, tracing, observability, and alerting.

- Hands-on experience with incident management and Root Cause Analysis (RCA).

- Understanding of infrastructure and application security, including secrets and vulnerability management.

- Experience working with Linux/Unix-based systems and production environments.

- Strong scripting/automation skills using languages such as Python, Bash, or similar.

Good to Have :

- Experience working with ML/AI infrastructure and LLM workloads.

- Experience with GPU infrastructure and model serving.

- Experience with Terraform and/or CloudFormation.

- Knowledge of MLOps, model deployment, and AI/ML platform engineering.

- Experience with cloud-native observability and distributed systems.

- Knowledge of security best practices for cloud and containerized environments.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...