HamburgerMenu
hirist

Lead DevOps Engineer - Cloud Infrastructure

GetHyr
6 - 12 Years
Noida

Posted on: 02/04/2026

Job Description

Description :


We are looking for a DevOps Lead who will own our cloud infrastructure, CI/CD pipelines, security, reliability, and scalability. You will work closely with engineering, product, and security teams to ensure our platforms are highly available, performant, and secure.


This is a hands-on leadership role where you will be responsible for designing, building, mentoring, and leading DevOps initiatives.


Key Responsibilities :


Infrastructure & Cloud :


- Own and manage cloud infrastructure (AWS preferred)


- Design highly available, scalable, and fault-tolerant systems


- Manage multiple environments (Production, Staging, QA, Dev)


- Implement Infrastructure as Code (Terraform / CloudFormation)


CI/CD & Automation :


- Design and maintain CI/CD pipelines (GitHub Actions / GitLab / Jenkins)


- Automate build, test, deployment, and rollback processes


- Improve deployment speed, reliability, and enable zero-downtime releases


Monitoring, Reliability & Security :


- Set up monitoring, alerting, and logging (Prometheus, Grafana, ELK, CloudWatch)


- Own uptime, SLAs, incident response, and post-mortems


- Implement security best practices (IAM, secrets management, vulnerability scans)


- Ensure compliance with security and audit requirements


Containerization & Orchestration :


- Manage Docker-based workloads


- Own Kubernetes (EKS preferred) setup, scaling, and upgrades


- Optimize resource utilization and cost efficiency


Leadership & Collaboration :


- Lead and mentor DevOps engineers


- Collaborate with backend, frontend, and QA teams


- Drive DevOps best practices across teams


- Participate in architecture and scalability discussions


AI & Modern Development :


- Experience integrating AI/LLM APIs (OpenAI, Anthropic, Hugging Face) into production systems


- Deploy and scale AI/ML workloads on cloud infrastructure (GPU instances, model serving)


- Manage rate limiting, cost optimization, and observability for LLM API usage


- Ensure security and compliance for AI-driven applications (data privacy, prompt injection, access control)


Required Skills & Experience :


Must-Have :


- 6+ years of experience in DevOps / SRE roles


- Strong experience with AWS


- Hands-on experience with Docker & Kubernetes


- Expertise in CI/CD pipelines


- Experience with Infrastructure as Code (Terraform preferred)


- Strong Linux and networking fundamentals


- Experience with monitoring, logging, and alerting tools


- Experience working on high-traffic SaaS or B2B platforms


- Experience integrating LLM APIs and deploying scalable AI/ML workloads


- Strong focus on cost optimization, observability, and security/compliance


Good to Have :


- Knowledge of security best practices and compliance


- Experience with cost optimization (FinOps)


- Exposure to multi-region deployments


- Scripting skills (Bash, Python, or Go)


What Were Looking For :


- Strong ownership mindset treats infrastructure like a product


- Excellent problem-solving and debugging skills


- Ability to balance speed, reliability, and security


- Clear communication and leadership skills


- Comfortable working in fast-paced, high-impact environments


info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...