HamburgerMenu
hirist

Platform Engineer - Kubernetes

Experis IT Private Limited
8 - 10 Years
Bangalore

Posted on: 17/06/2026

Job Description

We are seeking experienced Kubernetes Platform Engineers / Site Reliability Engineers (L2/L3) to support and maintain a large-scale cloud-native platform running on AWS and Kubernetes. The ideal candidate will have strong experience in Kubernetes operations, AWS infrastructure, Infrastructure as Code (Terraform), Python scripting, GitOps practices, and production incident management.

The role is primarily focused on platform support, troubleshooting, operational excellence, reliability engineering, and incident investigation rather than application development. Engineers will work closely with platform engineering, cloud operations, development, and support teams to ensure high availability, scalability, security, and performance of mission-critical services.

Key Responsibilities :

- Incident Hunting & Resolution : Perform deep-dive log analysis and trace execution paths across distributed systems to identify root causes of pipeline failures.

- Platform Support : Maintain the health of the Kratix orchestrator, ensuring containerized jobs and Python-based pipelines execute predictably across various clusters.

- Infrastructure as Code (IaC) : Manage and evolve AWS cloud resources using Terraform, ensuring consistency through GitOps workflows.

- Scaling & Optimization : Fine-tune cluster performance using Karpenter and Horizontal Pod Autoscalers (HPA) to manage fluctuating workloads efficiently.

- Policy & Governance : Monitor and troubleshoot cluster policies using Kyverno to ensure all workloads meet security and operational standards.

- Continuous Improvement : Identify patterns in recurring incidents and propose "self-healing" automation or Python scripts to reduce manual intervention (Toil).

Technical Qualifications :

- Kubernetes (Core) : Mastery of K8s primitives (Pods, Deployments, Services, CRDs, RBAC) and comfortable using CLI tools for live debugging.

- Python : Intermediate to Advanced proficiency, specifically for writing diagnostic scripts, interacting with APIs, and debugging containerized jobs.

- AWS Infrastructure : Hands-on experience with EKS, IAM roles, VPC networking, and S3.

- GitOps Ecosystem : Operational experience with Argo CD or Flux for continuous delivery and state reconciliation.

- Autoscaling Architecture : Proven ability to configure and troubleshoot Karpenter or Cluster Autoscaler to handle high-concurrency job execution.

- Policy Engines : Familiarity with Kyverno or OPA for managing cluster-wide guardrails.

Preferred Experience :

- Orchestration Logic : Prior exposure to platform orchestrators like Kratix, Crossplane, or Backstage.

- Log Management : Proficiency in ELK (Elasticsearch, Logstash, Kibana), Splunk, or Datadog for pattern recognition in high-volume log data.

- SRE Mindset : A "detective" approach to systemsvaluing the why of a failure as much as the how of the fix.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...