Posted on: 04/09/2026
About the Role:
We are looking for a skilled and proactive Site Reliability Engineer who will be responsible for ensuring the stability, availability, security, resiliency, and compliance of business-critical applications.
The role will involve production support, cloud infrastructure management, monitoring and observability, risk remediation, automation, and implementation of secure and resilient technology solutions.
Key Responsibilities:
- Ensure high availability, stability, reliability, and uptime of production applications.
- Monitor application and infrastructure health and proactively identify potential issues.
- Manage secrets, credentials, password rotations, certificates, and access controls.
- Perform vulnerability remediation, security patching, and technology risk mitigation.
- Manage certificate lifecycle activities including renewal, deployment, and expiry prevention.
- Perform software, application, operating system, and infrastructure upgrades.
- Identify and remediate technology obsolescence and unsupported components.
- Support technology resiliency, disaster recovery, high availability, and failover initiatives.
- Automate manual, repetitive, and operational processes using scripting and cloud-native technologies.
- Perform application and infrastructure cleanup and decommissioning of obsolete or unused resources.
- Provide L1/L2 production and on-call support, including incident investigation and resolution.
- Participate in Incident, Problem, and Change Management processes.
- Develop and maintain operational documentation, runbooks, and support procedures.
- Collaborate with global technical and non-technical stakeholders to deliver reliable technology solutions.
Technical Skills:
- AWS & Infrastructure: Strong hands-on experience with AWS (EC2, S3, IAM, Lambda, SQS, SNS, RDS, DynamoDB) and CloudFormation.
- Monitoring & Observability: Splunk, Grafana, AWS CloudWatch.
- Security & Risk Management: Secrets management, vulnerability assessment, patch remediation, certificate lifecycle management.
- DevOps / CI/CD: Git, GitHub Actions, Bamboo, ArgoCD.
- Automation: Python, Shell/Bash, or PowerShell.
- Production Support: Incident, Problem, and Change Management (L1/L2).
Domain Experience:
- Prior experience in Financial Services, Banking, Investment Banking, Asset Management, or FinTech is preferred.
- Experience working in highly regulated environments with strong security, risk, audit, and compliance requirements is highly desirable.
Ways of Working:
- Ability to work from the Gurugram office 3 days per week.
- Flexibility to work outside standard business hours when required for production support or critical activities.
- Experience working in an Agile delivery environment.
- Hands-on experience with JIRA and Confluence.
- Strong communication and stakeholder management skills.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1668588