Posted on: 07/07/2026
JD :
Operational Strategy & Stability :
- Incident Management : Take end-to-end ownership of production incidents, ensuring swift resolution and minimal customer impact for checkout and refund flows.
- System Reliability : Monitor the health of performant payment APIs and backend services, focusing on identifying trends before they become outages.
- RCA Excellence : Perform deep-dive root cause analysis on fuzzy system failures, translating complex logs into actionable engineering insights and permanent fixes.
Engineering Excellence :
- Automation : Eliminate "toil" by writing scripts and tools to automate repetitive manual tasks and operational workflows.
- Observability : Optimize our monitoring dashboards (Grafana/Kibana) to ensure anomalies in payment success rates (SR) are detected in real-time.
- Data Integrity : Work closely on investigating transactional discrepancies and reconciling massive data flows across internal ledgers and external bank partners.
Collaboration & Leadership :
- Liaise with Stakeholders : Act as the primary technical point of contact for Product, Business, and Customer Support teams for high-priority escalations.
- Knowledge Management : Maintain technical runbooks and documentation to ensure the highest industry standards for operational support within the team.
You should have :
- Educational Background : B.Tech or M.Tech in Computer Science, Information Technology, or a related technical discipline.
- Experience : 2+ years of experience in technical support, site reliability (SRE), or production support for large-scale distributed systems.
- Technical Proficiency : Strong expertise in SQL and working with complex database queries. Proficiency in at least one scripting language (Python, Shell, or Perl) for automation.
- Log Analysis : Deep experience with log aggregation and monitoring tools (e.g., ELK Stack, Splunk, Graylog, or New Relic).
- Analytical Mindset : Ability to evaluate production issues from multiple perspectivesbusiness, technical, and customerand translate them into logical solutions.
- Bias for Action : Ability to thrive in a high-pressure, fast-paced environment where quick decision-making is critical to maintaining platform uptime.
Good to Have :
- Prior experience in the Payments or Fintech domain (handling UPI, NetBanking, or Card transaction failures).
- Understanding of microservices architecture and containerization (Docker/Kubernetes).
- Familiarity with debugging Java or Scala application logs and understanding API.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
IT Management / IT Support
Job Code
1651965