HamburgerMenu
hirist

Site Reliability Engineer - Kubernetes/Prometheus

MiG Staffing
1 - 3 Years
Bangalore

Posted on: 05/09/2026

Job Description

About the Role :

We're looking for a Site Reliability Engineer to join our Site Reliability & Infrastructure Engineering team. We run entirely on the cloud, and this team owns production intelligence, reliability, and operational automation - moving from reactive incident response toward systems that observe, correlate, reason, and act. We're building the next generation of SRE : not dashboards and runbooks, but agents and automation that make our platform self-aware and self-healing at scale.

In this role, you'll work alongside Staff and Principal-level engineers who are setting that direction - learning how a real production platform is operated at scale, taking ownership of well-scoped pieces of it, and growing into bigger and more ambiguous problems as you go.

What You'll Do :

- Operate and support production Kubernetes infrastructure - assist with cluster upgrades, troubleshoot workload issues, and build a real understanding of how our platform runs at scale, with guidance from Staff and Principal engineers.

- Participate in on-call rotations, respond to incidents under supervision, and contribute to blameless postmortems - learning how to diagnose issues methodically and communicate clearly under pressure.

- Help build and maintain observability tooling - dashboards, alerts, and instrumentation across our OpenTelemetry/Prometheus/Grafana/Datadog stack - and start developing judgment for what's actually worth monitoring.

- Contribute to SLI/SLO definitions for the services you support, and help track error budgets day to day.

- Write and improve automation - scripts, tools, and small services that remove manual toil for the team.

- Support CI/CD pipeline reliability - help diagnose flaky pipelines and contribute to progressive delivery rollouts, learning how canary analysis and deployment safety checks work in practice.

- Get hands-on with our AI-assisted operational tooling - use, extend, and give feedback on the agents that senior engineers on the team are building for telemetry correlation and remediation.

What You'll Bring :

- Solid programming/scripting ability (Python, Go, or similar) and comfort working from the Linux command line.

- Some exposure to cloud platforms (AWS, GCP, or Azure) and containers/Kubernetes - coursework, an internship, or personal projects count.

- Curiosity about how distributed systems fail and get fixed, and genuine interest in learning production operations hands-on.

- 1 - 2 years of relevant experience (internship, new-grad role, or junior engineering position) in software engineering, DevOps, or a related field.

About You :

You're early in your career but already comfortable owning a well-defined problem end to end. You ask good questions, and you're not afraid to say "I don't know" and then go find out. You want to learn from people operating at Staff and Principal level, on a team where reliability isn't an afterthought - it's the actual work. You're excited about a team that's rethinking what SRE looks like as AI agents take on more of the routine investigation, and you want to help build that, not just read about it.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...