Posted on: 30/09/2026
Responsibilities :
- Design, build, and operate highly reliable cloud infrastructure supporting modern distributed applications. Develop automation and platform tooling using Python and Infrastructure as Code to improve engineering productivity and operational efficiency.
- Build and maintain scalable cloud infrastructure across AWS, GCP, or Azure using Terraform and modern cloud-native practices.
- Design and enhance observability solutions using Prometheus, Grafana, distributed tracing, and related monitoring technologies. Improve deployment workflows through CI/CD automation, GitOps practices, and infrastructure automation.
- Participate in incident response, troubleshoot production issues, and continuously improve platform reliability and resiliency.
- Build intelligent automation around monitoring, alerting, diagnostics, and operational workflows. Contribute to the design and development of next-generation observability and reliability products.
- Collaborate closely with the founding team on architecture decisions, platform design, and product development.
- Continuously improve scalability, reliability, security, and developer experience while working in a highly collaborative, ownership-driven engineering culture.
Requirements :
- Bachelor's degree in computer science, engineering, or a related field. 2 - 4 years of experience in platform engineering, cloud engineering, site reliability engineering, or infrastructure engineering.
- Strong hands-on experience with at least one cloud platform (AWS, GCP, or Azure) in production environments. Strong proficiency in Infrastructure as Code using Terraform.
- Good programming skills in Python or any modern programming language, with a focus on automation and tooling.
- Hands-on experience with Kubernetes and containerized workloads.
- Experience building automation for infrastructure provisioning, deployments, operational workflows, or platform tooling.
- Strong understanding of Linux systems, networking fundamentals, and cloud-native architectures. Experience working with observability and monitoring tools such as Prometheus, Grafana, Datadog, or similar.
- Good understanding of CI/CD pipelines, GitOps workflows, and infrastructure automation. Experience working on large-scale distributed systems and production infrastructure is preferred.
- Exposure to OpenTelemetry, security best practices, and open-source technologies is an added advantage. Excellent problem-solving, debugging, and communication skills, with the ability to work in a fast-paced startup environment.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1675920