Posted on: 09/09/2026
About the Role :
We are looking for an experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable, scalable, secure, and resilient technology platforms. The role will focus on improving system availability, performance, observability, automation, and operational efficiency across production environments.
Key Responsibilities :
- Design, implement, and maintain highly available, scalable, and resilient production infrastructure and services.
- Define and improve reliability standards, service-level objectives (SLOs), service-level indicators (SLIs), and error budgets.
- Monitor system health, application performance, infrastructure capacity, and availability using modern observability and monitoring tools.
- Build dashboards, alerts, and monitoring solutions to proactively identify performance degradation and potential production issues.
- Participate in incident response, troubleshoot complex production issues, and drive timely resolution of critical incidents.
- Conduct root cause analysis (RCA) for major incidents and implement preventive and corrective actions.
- Automate repetitive operational activities, deployments, infrastructure provisioning, and system maintenance.
- Develop and maintain CI/CD pipelines to enable reliable, secure, and efficient software releases.
- Work closely with development teams to improve application reliability, scalability, performance, and operational readiness.
- Implement Infrastructure as Code (IaC) using tools such as Terraform, Ansible, or equivalent technologies.
- Manage and optimize cloud infrastructure across services such as AWS, Azure, or GCP.
- Design and manage containerized workloads using Docker and Kubernetes.
- Implement effective backup, disaster recovery, high-availability, and business-continuity mechanisms.
- Identify infrastructure and application bottlenecks and recommend solutions to improve system performance and scalability.
- Establish capacity planning and resource utilization practices to support business growth.
- Contribute to security hardening, vulnerability remediation, access management, and secure infrastructure practices.
- Participate in production releases, change management, on-call support, and operational readiness activities.
- Maintain technical documentation, runbooks, architecture diagrams, operational procedures, and troubleshooting guides.
- Evaluate emerging technologies and recommend solutions that improve reliability, automation, observability, and operational efficiency.
- Mentor junior engineers and contribute to engineering best practices across the team.
Technical Skills :
- 7 - 10 years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, or related roles.
- Strong understanding of Linux/Unix systems, networking, system administration, and production environments.
- Hands-on experience with AWS, Azure, or GCP cloud platforms.
- Strong experience with Kubernetes and Docker/containerized environments.
- Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
- Strong knowledge of CI/CD tools and practices such as Jenkins, GitLab CI, GitHub Actions, Azure DevOps, or equivalent.
- Experience with monitoring and observability platforms such as Prometheus, Grafana, ELK/Elastic Stack, Splunk, Datadog, or New Relic.
- Strong scripting/programming skills in Python, Shell/Bash, Go, or similar languages.
- Good understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, load balancing, proxies, and firewalls.
- Experience with databases, caching systems, messaging platforms, and distributed applications.
- Strong understanding of microservices and distributed-system architecture.
- Experience implementing logging, metrics, tracing, alerting, and application performance monitoring.
- Good knowledge of Git and modern software development practices.
- Familiarity with cloud security, IAM, secrets management, vulnerability management, and secure configuration practices.
Reliability & Operations :
- Strong understanding of high availability, fault tolerance, scalability, resilience, and disaster recovery.
- Experience defining and tracking SLIs, SLOs, SLAs, and error budgets.
- Hands-on experience with incident management, escalation processes, post-incident reviews, and RCA.
- Ability to identify single points of failure and implement appropriate redundancy and failover mechanisms.
- Experience with capacity planning, performance optimization, and production readiness reviews.
- Ability to work effectively in an on-call and production-support environment.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1670129