Posted on: 01/05/2026
We are seeking a skilled Site Reliability Engineer (SRE) with strong development capabilities to ensure the reliability, scalability, and performance of our software systems.
This role combines software engineering and infrastructure expertise, focusing on building robust systems, automating operations, and improving overall system efficiency.
You will play a critical role in maintaining high system uptime, optimizing performance, and implementing automation to reduce manual intervention.
Key Responsibilities :
- Ensure high availability, reliability, and performance of production systems and services
- Design, build, and maintain scalable and resilient infrastructure
- Develop and implement automation tools and scripts to streamline operational tasks
- Monitor system health, troubleshoot issues, and perform root cause analysis for incidents
- Collaborate with development teams to improve system design, deployment processes, and observability
- Implement and manage CI/CD pipelines for efficient software delivery
- Manage and optimize cloud infrastructure and distributed systems
- Use Infrastructure as Code (IaC) practices to provision and manage environments
- Improve system performance through capacity planning, load testing, and tuning
- Establish best practices for monitoring, alerting, logging, and incident response
- Contribute to documentation and knowledge sharing across teams
Required Candidate Profile :
- Proven experience as a Site Reliability Engineer (SRE), DevOps Engineer, or similar role
- Strong proficiency in Python for scripting and automation
- Hands-on experience with automation tools such as Ansible, Puppet, Chef, or custom-built frameworks
- Solid understanding of Infrastructure as Code (IaC) tools like Terraform
- Experience working with cloud platforms (AWS, Azure, or Google Cloud)
- Strong knowledge of Linux/Unix systems, networking, and system administration
- Familiarity with containerization and orchestration tools (Docker, Kubernetes)
- Experience with monitoring and logging tools (Prometheus, Grafana, ELK stack, etc.)
- Good understanding of software development practices and full-stack environments
- Strong problem-solving and debugging skills
Preferred Qualifications :
- Experience in high-scale, distributed systems environments
- Knowledge of microservices architecture
- Exposure to security best practices and compliance standards
- Familiarity with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, etc.)
- Bachelors or Masters degree in Computer Science or a related field
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1632772