Posted on: 25/03/2026
Role Overview :
As an SRE DevOps Engineer, you will be at the forefront of ensuring the reliability, scalability, and performance of our critical systems and services. Your day-to-day will involve a blend of proactive monitoring, incident response, automation, and collaboration with development and operations teams.
You'll work closely with software engineers, product managers, and other stakeholders to build and maintain a robust infrastructure that supports our growing user base and business objectives. Your efforts will directly impact the availability and user experience of our platform, ensuring seamless service delivery and minimizing disruptions.
Key Responsibilities :
- Design, implement, and maintain CI/CD pipelines to automate software delivery and improve release velocity for engineering teams.
- Develop and maintain comprehensive monitoring and alerting systems to proactively identify and resolve potential issues before they impact users.
- Participate in incident response and post-mortem analysis to identify root causes and implement preventative measures to improve system resilience.
- Automate infrastructure provisioning and configuration management using Infrastructure-as-Code (IaC) principles to ensure consistency and repeatability.
- Collaborate with development teams to optimize application performance and scalability through code reviews, performance testing, and capacity planning.
- Contribute to the development and maintenance of our observability services to provide deep insights into system behavior and performance.
- Implement and enforce security best practices across our infrastructure and applications to protect sensitive data and prevent unauthorized access.
Required Skillset :
- Demonstrated ability to design, implement, and manage CI/CD pipelines using tools such as Jenkins, GitLab CI, or CircleCI.
- Proven expertise in building and maintaining observability solutions using tools like Prometheus, Grafana, ELK stack, or similar.
- Strong understanding of cloud computing platforms such as AWS, Azure, or GCP, and experience with containerization technologies like Docker and Kubernetes.
- Excellent problem-solving and troubleshooting skills, with the ability to analyze complex systems and identify root causes quickly.
- Solid understanding of Site Reliability Engineering (SRE) principles and practices, including monitoring, alerting, incident response, and automation.
- Effective communication and collaboration skills, with the ability to work effectively in a fast-paced, agile environment.
- Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1623639