HamburgerMenu
hirist

Lead DevOps & Site Reliability Engineer

TI Steps
4 - 8 Years
Multiple Locations

Posted on: 15/06/2026

Job Description

Job Description :

As a Lead DevOps & SRE, you will :

- Own the design, implementation, and continuous improvement of CI/CD pipelines across all products and services

- Architect and manage cloud infrastructure on AWS (or equivalent) including compute, networking, storage, databases, and managed services

- Define and enforce infrastructure-as-code practices using tools such as Terraform, CloudFormation, Ansible, or Pulumi

- Build and maintain observability, monitoring, alerting, and incident response systems to ensure platform reliability and uptime

- Establish and track SLIs, SLOs, and error budgets to drive data-informed reliability decisions

- Lead incident management, root cause analysis, post-mortems, and corrective action follow-through

- Implement security best practices across infrastructure, networking, secrets management, access controls, and compliance

- Manage containerized workloads using Docker and orchestration platforms such as Kubernetes or ECS

- Automate repetitive operational tasks, environment provisioning, scaling, backup, and disaster recovery

- Collaborate with backend and frontend engineers to ensure smooth deployments, rollback strategies, and zero-downtime releases

- Optimize infrastructure costs through rightsizing, reserved capacity planning, and usage monitoring

- Evaluate and introduce DevOps tooling, platforms, and practices that improve developer productivity and release velocity

- Contribute to AI-enabled product infrastructure by supporting GPU workloads, model serving, AI pipeline orchestration, and data platform needs where relevant

- Mentor engineers on DevOps practices, operational hygiene, and reliability culture

Why This Role Might Be for You :

- You want to build the infrastructure backbone that powers products creating real career and hiring impact

- Youre excited by the opportunity to design and scale cloud-native infrastructure for AI-powered platforms

- You enjoy solving complex operational problems and turning fragile systems into resilient, self-healing platforms

- You like staying hands-on with infrastructure while also guiding engineers and improving team operational maturity

- You want to work with a fast-moving team shaping products across career tech, hiring tech, AI, and workforce transformation

- Youre looking for a role where reliability, automation, security, speed, and engineering excellence all matter

Basic Qualifications :

- Bachelor's or master's degree in computer science, Engineering, Information Technology, or a related field

- 5- 8 years of experience in DevOps engineering, site reliability engineering, cloud infrastructure, or platform engineering

- Prior experience leading or owning infrastructure for production systems serving real users at scale

- Strong hands-on experience with cloud platforms such as AWS, GCP, or Azure

- Deep experience building and maintaining CI/CD pipelines using tools such as Jenkins, GitHub Actions, Bitbucket Pipelines, or ArgoCD

- Strong understanding of containerization (Docker) and container orchestration (Kubernetes, ECS, or similar)

- Experience with infrastructure-as-code tools such as Terraform, CloudFormation, Ansible, or Pulumi

- Strong understanding of networking, load balancing, DNS, SSL/TLS, firewalls, and VPC design

- Experience with monitoring, logging, and observability tools such as Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar

- Comfortable scripting and automating with Bash, Python, or Go

- Strong communication and collaboration skills with a practical problem-solving mindset

- Comfortable working in a fast-paced, evolving environment with ambiguity and changing priorities

- Understanding of how AI-enabled products influence infrastructure requirements, scaling patterns, and operational considerations

Preferred Qualifications :

- Experience with deploying products with technologies such as Python, Node.js, PostgreSQL, Redis, Nginx, or any similar application stacks

- Experience operating infrastructure for SaaS platforms

- Strong understanding of secrets management, IAM policies, network security, vulnerability scanning, and compliance frameworks

- Experience with database operations including backup, replication, failover, migration, and performance tuning for PostgreSQL, MongoDB, or similar

- Experience supporting AI/ML infrastructure including GPU instances, model serving (e.g. SageMaker, TorchServe), and data pipeline orchestration (e.g. Airflow, n8n)

- Experience with cost optimization strategies, FinOps practices, and cloud billing analysis

- Familiarity with tools such as Bitbucket, Jira, PagerDuty, Opsgenie, Vault, and collaboration platforms

- Interest in platform engineering, developer experience tooling, and internal developer platforms

- Curiosity about emerging trends in cloud-native architecture, AI infrastructure, and operational excellence

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...