Posted on: 11/05/2026
Role Overview :
As a Principal Site Reliability Engineer, you will be a key player in ensuring the reliability, performance, and scalability of our critical systems and infrastructure. You will collaborate closely with development, operations, and security teams to design, implement, and maintain resilient and highly available services. Your work will directly impact the experience of our customers by ensuring seamless and performant access to our platform.
Key Responsibilities :
- Design and implement scalable and resilient infrastructure solutions leveraging Kubernetes, Terraform, and cloud platforms (GCP, AWS, Azure).
- Develop and maintain monitoring and alerting systems using Prometheus and Grafana to proactively identify and resolve issues.
- Automate infrastructure provisioning, configuration management, and deployment processes to improve efficiency and reduce errors.
- Lead incident response efforts, performing root cause analysis and implementing preventative measures to minimize future disruptions.
- Collaborate with development teams to ensure that applications are designed for optimal performance, scalability, and reliability.
- Define and enforce service level objectives (SLOs) and service level agreements (SLAs) to ensure that our systems meet business requirements.
- Drive the adoption of DevOps best practices and a culture of continuous improvement across the organization.
- Mentor and guide junior SREs, fostering a collaborative and knowledge-sharing environment.
Required Skillset :
- Demonstrated ability to design, implement, and manage highly available and scalable infrastructure on cloud platforms (GCP, AWS, Azure).
- Proven expertise in containerization and orchestration technologies, particularly Kubernetes.
- Strong proficiency in infrastructure-as-code tools such as Terraform.
- Deep understanding of monitoring and alerting systems, including Prometheus and Grafana.
- Excellent scripting and automation skills using languages such as Python, Go, or Bash.
- Solid understanding of DevOps principles and practices, including CI/CD pipelines.
- Ability to effectively communicate technical concepts to both technical and non-technical audiences.
- Strong problem-solving and analytical skills, with a passion for identifying and resolving complex issues.
- Bachelor's degree in Computer Science or a related field.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1634869