Posted on: 29/05/2026
Description :
Were hiring a Senior Site Reliability Engineer (SRE) to strengthen the reliability, scalability, and security of our payments platforms, including payment gateways, clearing services, and core banking integrations. You'll sit at the intersection of software engineering and operations, using automation and engineering discipline to reduce toil, improve resilience, and raise observability standards across hybrid (onprem + cloud) environments.
Key responsibilities :
Reliability engineering & service ownership :
- Define and manage SLIs/SLOs and error budgets for critical payment services and supporting platforms.
- Improve availability, latency, and resilience through capacity planning, performance tuning, and reliability patterns (timeouts, retries, circuit breakers, bulkheads).
- Drive production readiness practices (runbooks, operational acceptance, resilience testing, DR readiness). Infrastructure & platform engineering (hybrid)
- Design and implement highly available infrastructure across onprem and cloud environments.
- Champion Infrastructure as Code (IaC) to standardize provisioning, configuration, and environment consistency.
- Support container platforms and middleware used in payments ecosystems (e.g., Kubernetes, MQ, application servers, databases).
Observability & incident management :
- Build and maintain monitoring, alerting, logging, and tracing to detect transaction-flow anomalies early.
- Lead or support major incident response as a technical incident commander for payment-impacting events.
- Run blameless postmortems, identify systemic fixes, and track actions through to completion.
- Participate in a tiered on-call rotation supporting 24x7x365 operations.
Automation & CI/CD :
- Reduce operational toil through automation (self-healing, auto-remediation, standardized runbooks).
- Improve CI/CD pipelines to enable safe, secure, and repeatable deployments (including controls appropriate for regulated environments).
- Partner with engineering teams to embed reliability into delivery (testing, release strategies, rollback plans). Payments domain focus
- Troubleshoot issues across the payment lifecycle: routing validation clearing settlement reconciliation.
- Apply knowledge of transaction integrity (idempotency, ordering, consistency) in distributed payment
systems.
- Support adoption and reliability of payment rails and standards (e.g, ISO 20022, ACH/ & Real-time payments).
Required experience (7- 10 years) :
- 7- 10 years in SRE / DevOps / Production Engineering / Systems Engineering roles supporting mission-critical services.
- Demonstrable experience operating payments or banking platforms (banking/FinTech domain exposure expected; not necessarily 15+ years).
- Experience working in highly regulated environments with strong security and audit expectations.
Technical skills :
- Programming/Scripting : Go, Python, Java, and/or Shell (automation, tooling, reliability improvements).
- Containers & orchestration : Kubernetes, Docker (deployment patterns, scaling, reliability).
- Cloud & hybrid : AWS and/or GCP (plus strong onprem integration experience).
- Middleware & data platforms (nice to have / relevant) : IBM WebSphere (WAS), IBM MQ, Oracle, DB2,
MySQL.
- IaC & config management : Terraform, Ansible, CloudFormation.
- Observability : Datadog / AppDynamics / Dynatrace, Splunk, Prometheus, Grafana (metrics, logs, traces).
- CI/CD : Jenkins, GitLab CI, GitHub Actions (secure pipelines, release controls).
Education :
- UG : B.Tech / B.E. in Any Specialization
- PG : M.Tech in Any Specialization, MCA in Any Specialization
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1640196