HamburgerMenu
hirist

SatSure - Senior Cloud/DevOps Engineer

hirist.tech
5 - 10 Years
Bangalore

Posted on: 11/09/2026

Job Description

Note : If screened-in, you will be invited for initial rounds on 10th October 2026 (Saturday) in Bangalore.


Role : Senior Cloud/DevOps Engineer (SDE - 3)


About SatSure :

SatSure is a deep-tech decision intelligence company operating at the nexus of agriculture, infrastructure, and climate action. We turn earth observation data into actionable insights for governments, financial institutions, and enterprises across the developing world - at scale, with reliability.

Our platform team owns the infrastructure backbone that powers SatSure's AI/ML products: multi-cloud Kubernetes clusters, LLM inference pipelines, geospatial data platforms, and the internal developer tooling used by every engineering team.

About the Role :

We are looking for a Senior DevOps & MLOps Engineer to join our Platform & DevOps team. You will design, build, and operate cloud-native infrastructure that supports ML model serving, data pipelines, and developer platforms across AWS, GCP, and Azure.

Roles & Responsibilities :

ML Platform & LLM Infrastructure :

- Own and operate Kubernetes-based ML platform on EKS - supporting LLM inference (KServe), distributed compute (Dask/Ray), and workflow orchestration (Apache Airflow).

- Partner with data science and ML teams to design, deploy, and scale ML workloads - including GPU scheduling, autoscaling, resource isolation, and SLO-driven reliability.

- Architect, deploy, and optimize Ray clusters on Kubernetes for distributed ML workloads.

Multi-Cloud Platform & Infrastructure :

- Design, build, and maintain cloud-native infrastructure across AWS (primary), GCP, and Azure - using Kubernetes (EKS / GKE / AKS), Terraform, Helm, and ArgoCD.

- Drive GitOps adoption and platform standardization - define reusable infrastructure patterns, Helm charts, and deployment workflows.

- Manage Kubernetes platform operations - cluster lifecycle, Karpenter-based autoscaling, multi-tenancy, and workload isolation.

- Implement and maintain service mesh (Istio) - mTLS enforcement, traffic policies, and observability.

- Maintain and improve the internal developer platform (Backstage IDP).

Observability & Reliability Engineering :

- Build and maintain full-stack observability infrastructure - metrics (Prometheus / Mimir), logs (Loki), traces (Tempo), and dashboards (Grafana) integrated with OpenTelemetry.

- Define SLIs, SLOs, and error budget policies for production ML and platform services; lead incident response and post-mortem reviews.

FinOps & Cost Engineering :

- Implement Kubernetes cost attribution and chargeback using Kubecost / OpenCost.

- Continuously optimize cloud spend through workload right-sizing, spot/preemptible usage, and resource scheduling strategies.

Platform Security & Governance :

- Manage AWS multi-account governance using Control Tower, SCPs, GuardDuty, and IAM Identity Center.

- Own OIDC identity and SSO infrastructure integrated across internal tooling.

- Support compliance and audit processes - ISO 27001, CIS Benchmarks, Well-Architected Reviews, and VAPT assessments.

Requirements :

Must Have :

- 5+ years of hands-on platform, DevOps, or SRE experience in production environments.

- Strong Kubernetes expertise - cluster operations, Helm, RBAC, autoscaling (Karpenter / Cluster Autoscaler), multi-tenancy; EKS experience preferred.

- Infrastructure as Code - Terraform (advanced), Ansible.

- AWS expertise - EC2, EKS, S3, RDS, IAM, VPC, CloudWatch, Control Tower, GuardDuty.

- GitOps & CI/CD - ArgoCD, Bitbucket Pipelines / Jenkins.

- Observability - hands-on with Prometheus, Grafana, and at least one of: Loki, Tempo, OpenTelemetry, Datadog, or ELK.

- Scripting & automation - Python and Bash.

Why SatSure :

- Real Production Scale : LLM inference, geospatial data pipelines, and multi-cloud Kubernetes.

- High Ownership : You architect systems end-to-end.

- Meaningful Impact : Your infrastructure powers products used by governments and institutions across the developing world.

- Growth & Benefits : Learning allowances, broadband, medical insurance, best-in-class leave policy, and hybrid work from Bengaluru.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...