Posted on: 10/10/2026
Job Purpose :
Analysing, troubleshooting, and designing vital services, platforms, and infrastructure on GCP while always thinking about reliability, scalability, resilience, security, and performance.
Job Responsibilities :
- Help build a Site Reliability Engineering culture by sharing the best practices, approaches, documentation, and code with other engineering teams.
- Apply automation and software to any tasks or parts of the system which are performed manually.
- Able to troubleshoot complicated, cross platform issues handling OS, Networking, Database in a cloud-based SaaS environment and handle live production incidents.
- Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.
- Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.
Actionable :
- Design, write, ship, and motivate the creation of software and systems to increase observability, product reliability, and organizational efficiency.
- Maintain and monitor deployment, orchestration of the servers, docker containers, databases, and general backend infrastructure.
- Develop Run Books/Standard Operating Procedure for recurring Production issues, also working on a permanent solve.
- Perform Incident Analysis on a regular basis with the intention of preventing and finding a long-term solve for Incidents.
Tech Stack :
- GCP, Python Scripting
- Monitoring and analyzing infrastructure performance using standard performance monitoring tools
- Containerization : Docker and orchestration (Kubernetes)
- Infrastructure As Code : Terraform, Cloud Formation, Ansible
- Large-scale databases and distributed technologies : Kafka and Confluent Platform Kafka
- Basic programming and scripting skills
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1677782