HamburgerMenu
hirist

KnowledgeWorks Global - Senior DevOps Engineer - Docker/Kubernetes

Knowledgeworks Global
6 - 12 Years
Mumbai

Posted on: 30/07/2026

Job Description

Sr. DevOps Engineer

Observability Strategy (Web & AI Applications) :

- Design and implement an end-to-end observability & alert stack covering both traditional web applications and AI/ML services.

- Develop and maintain tooling such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalent.

- Build AI-specific observability : model latency/throughput tracking, GPU utilization, token usage, drift detection, and inference quality signals.

Infrastructure Scaling & Reliability :

- Design and manage infrastructure capable of scaling across on-premise data centers and public cloud.

- Own capacity planning, load testing, auto-scaling, and cost-optimization initiatives across compute, storage, and networking.

- Implement Infrastructure as Code (Terraform, Ansible, or equivalent) to ensure environments are reproducible, version-controlled, and auditable.

- Lead disaster recovery, backup, and high-availability strategy for critical systems.

DevOps & CI/CD Delivery :

- Partner closely with Solution Architects to translate project and system designs into concrete DevOps execution plans.

- Design, build, and maintain CI/CD pipelines (Jenkins, GitHub Actions or ArgoCD/Flux for GitOps) across multiple projects and teams.

- Containerize and orchestrate applications using Docker and Kubernetes, including Helm chart and manifest management.

- Embed security and compliance checks (SAST/DAST, secrets scanning, image scanning) directly into the delivery pipeline (DevSecOps).

GPU Infrastructure & Model Deployment :

- Provision, configure, and manage GPU infrastructure (on-prem clusters and cloud GPU instances) for model training and inference.

- Deploy, scale, and monitor ML/LLM models in production using tools such as Triton Inference Server, vLLM, or similar.

- Optimize GPU utilization, cost, and throughput across multi-tenant workloads; manage CUDA/driver/toolkit versions.

- Collaborate with data science/ML engineering teams on MLOps pipelines model versioning, experiment tracking, and model registry.

Requirements :

- 5-10 years of hands-on DevOps/SRE/Infrastructure engineering experience, including at least 23 years in a senior or lead capacity.

- Deep expertise in Git-based workflows and repository management at scale (GitHub/GitLab/Bitbucket).

- Proven experience designing observability stacks (Prometheus, Grafana, ELK/EFK, Datadog, New Relic, OpenTelemetry).

- Strong background in cloud platforms (AWS, Azure, and/or GCP) and on-premise/hybrid infrastructure.

- Expert-level skills with Infrastructure as Code (Terraform, Ansible, CloudFormation, or Pulumi).

- Strong Kubernetes and Docker experience, including multi-cluster and multi-environment management.

- Hands-on experience building and maintaining CI/CD pipelines end to end.

- Working knowledge of GPU infrastructure (NVIDIA CUDA, drivers, NCCL) and experience deploying ML/AI models to production.

- Proficiency in scripting/automation languages : Python, Bash, and/or Go.

- Solid understanding of networking, load balancing, DNS, and security fundamentals in distributed systems.

- Experience partnering with architects and engineering leads to translate designs into infrastructure and delivery plans.

- Hands-on with SAST, DAST, and SCA tooling (e.g., SonarQube, Snyk, Checkmarx, OWASP DependencyCheck) integrated directly into CI/CD pipelines.

- Container and image security : vulnerability scanning (Trivy, Grype, Clair), minimal/hardened base images, and signed/verified image provenance (Cosign/Sigstore).

- Familiarity with Secrets management and credential hygiene using tools such as HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault.

- Excellent communication skills and comfort operating cross-functionally with development, data science, and product teams.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...