HamburgerMenu
hirist

Data/Site Reliability Engineer - Native Cloud Infrastructure

svan global consultancy
4 - 6 Years
Bangalore

Posted on: 24/06/2026

Job Description

About the Role :

We are looking for an experienced Infra Data Engineer / Data SRE to build, operate, and scale modern data infrastructure platforms with a strong focus on Kubernetes, reliability, automation, scalability, and cost optimization.

In this role, you will work closely with Data Engineering and Platform teams to manage distributed data systems, improve operational efficiency, and build reliable cloud-native infrastructure platforms.

Key Responsibilities :

Kubernetes & Platform Operations :

- Manage and optimize Kubernetes-based data infrastructure platforms in production environments.

- Support cluster upgrades, autoscaling, workload scheduling, and resource optimization.

- Implement Kubernetes best practices including StatefulSets, Persistent Volumes, Pod Disruption Budgets, and affinity rules

- Ensure platform reliability, high availability, and disaster recovery readiness.

Data Platform Engineering & Reliability :

- Operate and support distributed systems such as Kafka, Spark, and Airflow.

- Build and maintain observability solutions using Prometheus, Grafana, and centralized logging platforms.

- Participate in incident management, root cause analysis, operational automation, and on-call support.

- Create and maintain operational runbooks, dashboards, and platform documentation.

Automation, CI/CD & Cost Optimization :

- Build and maintain CI/CD and GitOps workflows for infrastructure and platform deployments.

- Automate infrastructure provisioning using Terraform, Helm, and related IaC tools.

- Monitor and optimize infrastructure utilization across compute, storage, and networking.

- Implement security best practices including RBAC, secrets management, and vulnerability scanning.

Required Skills & Qualifications :

- 4-5 years of experience in Infrastructure Engineering, DevOps, Platform Engineering.

- Strong hands-on experience with Kubernetes in production environments.

- Experience operating distributed systems such as Kafka, Spark, or Airflow.

- Experience with CI/CD pipelines, GitOps, Terraform, and Helm.

- Familiarity with observability tools such as Prometheus, Grafana, and ELK/OpenSearch.

- Good understanding of Linux systems, networking, and cloud infrastructure.

- Scripting/programming experience in Python, Bash, or Go.

- Strong troubleshooting and incident management skills.

Good to Have :

- Experience with autoscaling tools such as HPA, VPA, or KEDA.

- Exposure to cloud platforms such as GCP or AWS.

- Experience with Vault, Trivy, or similar security and secrets management tools.

- Understanding of cost optimization and FinOps practices.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...