Posted on: 11/08/2026
Role Overview:
Hiring VP - Service Operations for one of our clients.
Service Reliability & Availability:
- Deliver Five Nines Availability (99.999%): Build and enforce architectural and operational practices that ensure global transit payment systems achieve and sustain ultra-high uptime. This includes proactive monitoring, high-availability design enforcement, and automated failover systems across cloud and on-premises platforms.
- Govern Service Level Objectives (SLOs) & Error Budgets: Define, track, and report SLOs for all critical services, ensuring error budgets are managed responsibly to balance reliability with change velocity.
- Transaction Performance Management: Guarantee low-latency, high-throughout processing across all environments, actively tuning Oracle databases, Kubernetes clusters, UCS fabrics and cloud services for peak demand conditions.
- Preventative Maintenance Program: Own a proactive, structured maintenance strategy (firmware updates, DB patching, load balancing, failover readiness) to ensure reliability and reduce risk of unplanned outages.
Incident, Problem & Change Management:
- Global Incident Oversight: Lead the 24x7 incident response process across UK and India operations centers, ensuring ?95% SLA compliance for response and resolution.
- Root Cause & Blameless Postmortems: Mandate blameless postmortems for every significant incident, driving systemic fixes to prevent recurrence and sharing lessons learned across all regions.
- Problem Management: Establish a structured problem management process to identify trends, reduce repeat incidents, and address root technical or process flaws.
- Change Governance: Oversee change management to balance service stability with innovation. Implement automated pipelines where possible to reduce manual error and ensure changes are tested for resilience before release.
- Recovery Metrics: Continuously track and drive down Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR) with quarterly improvement goals.
Observability, Automation & SRE Practices:
- Full-Stack Observability: Deploy and govern monitoring solutions across metrics, logs, and traces, enabling real-time visibility into system health and early anomaly detection.
- Automation of Toil: Drive a culture of automation by eliminating repetitive manual tasks in patching, scaling, failover, and incident response. Track progress with explicit automation coverage targets.
- Chaos Engineering & Resilience Testing: Institutionalize chaos testing and DR drills (game days) to validate RTO/RPO readiness and system recovery under stress conditions.
- Capacity & Performance Engineering: Oversee predictive capacity planning to ensure the system scales automatically to meet load and maintain performance even under extreme demand conditions.
- Proactive Reliability Improvements: Fund and prioritize engineering initiatives aimed at improving long-term service resilience, not just reactive firefighting.
Global Operations & Customer Support:
- Global Ops Center Leadership: Direct two 24x7 global operations centers (UK and India) as the backbone of global service delivery. Ensure staffing, shift rotations, and runbooks meet the highest standards of responsiveness and reliability.
- Localized Team Oversight: Manage in-country support teams that ensure local compliance, and customer-specific responsiveness.
- Customer Contact Centers: Own the operations of customer-facing contact centers, ensuring tight integration with back-end service teams to provide consistent, rapid, and high-quality customer experience.
- Standardization Across Regions: Drive consistency of ITIL-aligned service management processes globally, ensuring customers experience the same high standards regardless of geography.
Compliance, Risk & Audit Readiness:
- Global Compliance Ownership: Ensure operations remain compliant with ISO27001, PCI DSS v4.1, Essential Eight, Cyber Essentials, GDPR, and all other local privacy regulations in customer jurisdictions.
- Audit Readiness & Evidence: Maintain continuous audit readiness with robust evidence trails across all systems, processes, and controls, avoiding last-minute remediation before audits.
- Corporate CISO & Service Assurance Partnership: Work in lockstep with the Corporate CISO and Service Assurance functions to embed security controls, risk frameworks, and governance processes into day-to-day operations, aiming for regulatory compliance excellence.
- Risk Reporting: Provide transparent reporting to the COO and Senior Leadership Teams on compliance posture, vulnerabilities, audit results, and remediation progress.
- Vulnerability Management: Ensure vulnerabilities are tracked, prioritized, and remediated on schedule, with operational accountability for timely fixes.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Senior Management
Job Code
1662276