Posted on: 04/07/2026
We are looking for an experienced Lead Site Reliability Engineer (SRE) / Resiliency Engineer to help define and implement resiliency standards, improve system reliability, automate infrastructure and application management, and drive operational excellence across engineering teams.
Key Responsibilities :
- Architect and manage highly available, scalable infrastructure on cloud platforms to ensure seamless service delivery for global clients.
- Lead the design and implementation of container orchestration strategies using Kubernetes to improve deployment velocity and system stability.
- Drive the adoption of Infrastructure as Code (IaC) using Terraform to standardize environment provisioning and reduce configuration drift.
- Develop and maintain automation scripts using Python to eliminate operational toil and enhance the efficiency of incident response processes.
Mandatory Skills :
- 6 - 10 years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Engineering.
- Strong understanding of distributed systems, networking, and modern application architectures.
- Experience designing and implementing highly available, resilient, and fault-tolerant systems.
- Hands-on experience with cloud and on-premises infrastructure.
- Experience with observability tools, including metrics, logging, tracing, dashboards, and alerting.
- Proficiency in at least one programming or scripting language such as Python, Java, or C#.
- Hands-on experience with Infrastructure as Code (IaC) or Configuration Management tools such as Terraform, Bicep/ARM, Chef, Puppet, or Ansible.
- Experience implementing Infrastructure as Code and deployment automation using CI/CD pipelines.
- Strong analytical, problem-solving, and communication skills.
Did you find something suspicious?
Posted by
Vishal Harsule
TA Business Partner at iVector
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1651360