HamburgerMenu
hirist

Job Description

We are looking for an experienced Lead Site Reliability Engineer (SRE) / Resiliency Engineer to help define and implement resiliency standards, improve system reliability, automate infrastructure and application management, and drive operational excellence across engineering teams.


Key Responsibilities :


- Architect and manage highly available, scalable infrastructure on cloud platforms to ensure seamless service delivery for global clients.


- Lead the design and implementation of container orchestration strategies using Kubernetes to improve deployment velocity and system stability.


- Drive the adoption of Infrastructure as Code (IaC) using Terraform to standardize environment provisioning and reduce configuration drift.


- Develop and maintain automation scripts using Python to eliminate operational toil and enhance the efficiency of incident response processes.


Mandatory Skills :

- 6 - 10 years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Engineering.

- Strong understanding of distributed systems, networking, and modern application architectures.

- Experience designing and implementing highly available, resilient, and fault-tolerant systems.

- Hands-on experience with cloud and on-premises infrastructure.

- Experience with observability tools, including metrics, logging, tracing, dashboards, and alerting.

- Proficiency in at least one programming or scripting language such as Python, Java, or C#.

- Hands-on experience with Infrastructure as Code (IaC) or Configuration Management tools such as Terraform, Bicep/ARM, Chef, Puppet, or Ansible.

- Experience implementing Infrastructure as Code and deployment automation using CI/CD pipelines.

- Strong analytical, problem-solving, and communication skills.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...