HamburgerMenu
hirist

Site Reliability Engineer - Cloud Infrastructure

Scaling Theory Technologies
2 - 4 Years
Bangalore

Posted on: 30/09/2026

Job Description

Responsibilities :

- Design, build, and operate highly reliable cloud infrastructure supporting modern distributed applications. Develop automation and platform tooling using Python and Infrastructure as Code to improve engineering productivity and operational efficiency.

- Build and maintain scalable cloud infrastructure across AWS, GCP, or Azure using Terraform and modern cloud-native practices.

- Design and enhance observability solutions using Prometheus, Grafana, distributed tracing, and related monitoring technologies. Improve deployment workflows through CI/CD automation, GitOps practices, and infrastructure automation.

- Participate in incident response, troubleshoot production issues, and continuously improve platform reliability and resiliency.

- Build intelligent automation around monitoring, alerting, diagnostics, and operational workflows. Contribute to the design and development of next-generation observability and reliability products.

- Collaborate closely with the founding team on architecture decisions, platform design, and product development.

- Continuously improve scalability, reliability, security, and developer experience while working in a highly collaborative, ownership-driven engineering culture.

Requirements :

- Bachelor's degree in computer science, engineering, or a related field. 2 - 4 years of experience in platform engineering, cloud engineering, site reliability engineering, or infrastructure engineering.

- Strong hands-on experience with at least one cloud platform (AWS, GCP, or Azure) in production environments. Strong proficiency in Infrastructure as Code using Terraform.

- Good programming skills in Python or any modern programming language, with a focus on automation and tooling.

- Hands-on experience with Kubernetes and containerized workloads.

- Experience building automation for infrastructure provisioning, deployments, operational workflows, or platform tooling.

- Strong understanding of Linux systems, networking fundamentals, and cloud-native architectures. Experience working with observability and monitoring tools such as Prometheus, Grafana, Datadog, or similar.

- Good understanding of CI/CD pipelines, GitOps workflows, and infrastructure automation. Experience working on large-scale distributed systems and production infrastructure is preferred.

- Exposure to OpenTelemetry, security best practices, and open-source technologies is an added advantage. Excellent problem-solving, debugging, and communication skills, with the ability to work in a fast-paced startup environment.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...