Job Description :
Key Responsibilities :
- Architect and implement robust, self-healing cloud infrastructure on AWS to ensure maximum uptime and system resilience for our global client base.
- Define and enforce rigorous SLI/SLO/SLA frameworks to provide clear visibility into system health and drive data-backed improvements in service reliability.
- Lead large-scale automation initiatives to eliminate manual toil, thereby increasing deployment velocity and reducing human error in production environments.
- Establish advanced observability practices, including distributed tracing and proactive alerting, to minimize mean time to detection (MTTD) and resolution (MTTR) for complex incidents.
Must have :
- Demonstrated experience with the observability tools, automation and resilience and cloud architecture.
Nice to have :
- Demonstrated experience or knowledge with the following Observability Tools : Splunk, Honeycomb, CloudWatch, Dynatrace, AppDynamics.
- Metrics & Dashboards : SLI/SLO/SLA generation, JIRA dashboards, SLO trackers.
- Monitoring & Alerting : CASE methodology, anomaly detection, predictive alerting, synthetic monitoring.
- Automation & Resilience : Blue Prism, Python, UiPath, Chaos Engineering, FMEA.
- Cloud Architecture : AWS, Azure.
- Understanding of Operating Systems and Networks.
- Leadership & Influence Coaching & Enablement : Support product teams in elevating operational maturity and reliability readiness.
- Strategic Thinking : Align observability goals with business outcomes and OKRs.
- Community Engagement : Contribute to SRE guilds, CoPs, and internal training initiatives.
- Influence Direction : Work with teams to prioritize and implement resiliency initiatives.
- Incident Management : Support teams with debugging and problem solving during Major Incidents.