HamburgerMenu
hirist

Bajaj Life Insurance - Chief Manager - Application Management/Service Reliability & Observability

BAJAJ LIFE INSURANCE
8 - 12 Years
Pune

Posted on: 04/08/2026

Job Description

Role : Bajaj Life Insurance - Chief Manager - Application Management

Experience : 8 to 12 years

Role Title : Chief Manager Application Management (Service Reliability & Observability)

Reports To : Vice President Service Reliability & Observability

Company : BAJAJ LIFE Insurance

Function/ Department : Technology

JOB PURPOSE :

This techno functional role is responsible for ensuring the availability, stability, and performance of business-critical insurance applications (policy administration, claims, payouts, renewals, and policy servicing platforms) through effective production support, DevOps practices, and team leadership. The role owns end-to-end incident management, release and change management, and continuous improvement of application reliability, while building and mentoring a team of junior production support engineers to deliver consistent, SLA-compliant service to business and policyholders. Any Life Insurance Background would be helpful.

PRINCIPAL ACCOUNTABILITIES :

Application Availability & Reliability Management :

- Own end-to-end availability of policy administration, claims, and servicing applications against defined SLA/OLA targets (e.g. 99.8%+ uptime).

- Implement and continuously improve monitoring, alerting, and observability (APM, logs, synthetic checks) to detect issues proactively before business impact.

- Analyze recurring incidents to identify and remediate root causes, driving down repeat failures and reducing Mean Time to Detect (MTTD) and Mean Time to Restore (MTTR).

- Drive capacity planning and performance tuning (application, database, infrastructure) to prevent availability degradation during peak policy-servicing cycles (e.g. renewal season, month-end/quarter-end).

- Own and periodically test disaster recovery (DR) and business continuity procedures for critical applications, ensuring rapid restoration with minimal data loss.

Production & Product Support Management :

- Ensure timely triage, prioritization, and resolution of production incidents and service requests as per the P1-P4 severity framework and agreed SLA/OLA timelines.

- Act as escalation point for critical (P1/P2) incidents, coordinating war-room bridges, vendor/IT teams, and communicating status to stakeholders until closure.

- Institutionalize a Root Cause Analysis (RCA) process for all major incidents, ensuring corrective and preventive actions (CAPA) are tracked to closure.

- Oversee ticket/queue management (via ITSM tools such as JIRA/ServiceNow) to ensure ageing, pendency, and SLA breaches are actively managed and reported.

- Manage patching, version upgrades, and security vulnerability remediation across application environments without disrupting business operations.

Policy Servicing Functional & Technical Issue Resolution :

- Ensure functional and technical issues affecting policy servicing journeys New Business, Renewals, Endorsements, Claims, Payouts, Free-Look Cancellation, Policy Loan, Grievance, and Revival are resolved accurately and within SLA.

- Coordinate with business teams (Operations, Claims, Customer Service) to understand servicing-related pain points and translate them into technical fixes or enhancements.

- Validate that fixes/releases affecting policy servicing modules are tested (functional + regression) before go-live to avoid recurrence of customer-impacting defects.

- Track and report on policy-servicing-related technical issues (SDC%, TAT%, NOP, pendency) to demonstrate service quality improvements over time.

Change, Release & DevOps Process Management :

- Own the change management and release process for production deployments, ensuring controlled, low-risk releases with rollback plans.

- Champion DevOps and CI/CD practices (build/deploy automation, infrastructure-as-code, environment consistency) to improve release velocity and quality.

- Verify completion of development, testing, and approvals prior to production release; ensure post-deployment monitoring and sign-off.

- Assess and manage the impact of proposed changes on live applications, maintaining a change control log and risk register.

- Manage API integrations and data flows between core applications (policy admin, payment gateway, eKYC, IGMS, etc.) to ensure end-to-end process integrity.

Stakeholder & Vendor Management :

- Establish and run governance cadences (daily stand-ups, weekly/monthly reviews) with business and technical stakeholders on incidents, changes, and service health.

- Prepare and present Weekly/Monthly Senior Management Communication covering availability, incident trends, SLA compliance, and improvement actions.

- Liaise with application vendors/system integrators (e.g. core policy administration platform partners) on defect fixes, patches, and enhancement delivery.

- Coordinate with Business Analysts, Project Managers, Infosec, and Compliance teams to ensure production changes meet regulatory (IRDAI) and internal audit requirements.

- Maintain transparent, timely communication with stakeholders during incident resolution to avoid information gaps.

Team Leadership & Capability Building :

- Lead, mentor, and manage day-to-day performance of a team of junior production support engineers, including shift/roster planning for 24x7 coverage where applicable.

- Define individual and team KPIs aligned to availability, TAT, and quality goals; conduct regular performance reviews and feedback.

- Build team capability through structured upskilling on ITIL practices, monitoring/observability tools, cloud platforms, scripting/automation, and insurance domain knowledge.

- Ensure adequate documentation, knowledge transfer, and runbooks exist so that support quality does not depend on individual tribal knowledge.

- Foster a culture of ownership, proactive problem-solving, and continuous improvement within the support team.

MAJOR CHALLENGES :

Balancing the need for strong, consistent governance rhythms (SLA/OLA reviews, RCA closures, senior management reporting) against the team's operational bandwidth, particularly during high-load periods such as month-end/quarter-end policy servicing cycles. Correlating user experience issues with underlying system logs, alerts, and infrastructure signals to derive a precise, prioritized action plan for application stability improvements, while managing a lean team and multiple concurrent production issues.

DECISIONS :

- Prioritization and severity classification (P1-P4) of incoming production and support issues.

- Go/no-go decisions on production releases and emergency changes, within defined change-control authority.

- Escalation timing and routing for major incidents to senior stakeholders and vendors.

- Documenting, tracking, and prioritizing RCA-driven technical improvement actions and change requests.

- Allocation of team workload/shift coverage across incidents, changes, and BAU support activities.

INTERACTIONS :

Internal Clients :

- IT Leadership, Application Development Teams, Infrastructure/Cloud Teams, Business Operations (New Business, Claims, Servicing), Business Analysts, Project Managers, Information Security, Compliance/Audit, Customer Service.

External Clients :

- Application/Platform Vendors and System Integrators, Cloud Service Providers, Third-Party API Partners (Payment Gateway, eKYC, IGMS), External Auditors/Regulatory Bodies as required.

DIMENSIONS :

Financial Dimensions (FY 24) : NA.

Other Dimensions (FY 24) : Total Team Size : 2/3, Number of Direct Reports : 2.

SKILLS AND KNOWLEDGE :

Educational Qualifications :

- Bachelor's degree in computer science, information technology, or a related field (or equivalent work experience).

- Proven experience in log analysis, DB performance fine-tuning measures, application support and management roles.

Work Experience :

- 8+ years of overall IT experience, including at least 3-4 years in a production support, application management, or DevOps leadership role.

- Proven experience managing a team, ideally within BFSI/insurance or another regulated industry.

- Track record of improving application availability, reducing MTTR, and driving incident/problem management maturity.

- Experience with core insurance systems (policy administration platforms, claims systems) preferred.

Technical Skills :

- Strong knowledge of ITIL processes : Incident, Problem, Change, and Release Management.


- Hands-on exposure to monitoring/observability tools (APM, log management, dashboards) and ITSM platforms (JIRA, ServiceNow, or equivalent).


- Working knowledge of CI/CD pipelines, scripting/automation, containerization (Kubernetes/Docker), and cloud platforms (GCP/AWS/Azure).


- Familiarity with database performance tuning (Oracle/PostgreSQL) and application log analysis. Understanding of API-based integrations and microservices architectures.

Domain & Regulatory Knowledge :

- Understanding of Indian life insurance business processes New Business, Renewals, Claims, Payouts, Endorsements, Free-Look, Revival, Grievance.


- Awareness of IRDAI regulatory requirements and data protection norms (DPDP Act 2023) as applicable to production systems and customer data.

Behavioral Competencies :

- Strong analytical and problem-solving skills with a structured, root-cause-driven mindset.


- Excellent stakeholder communication skills, including experience presenting to senior/CXO audiences.


- Ability to lead and develop a team under pressure, balancing governance rigor with operational bandwidth.


- High ownership and accountability for outcomes, with a continuous-improvement orientation.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...