HamburgerMenu
hirist

Senior DevOps Engineer - AWS Infrastructure

Pravi HR Advisory
5 - 8 Years
Mumbai

Posted on: 08/06/2026

Job Description

Role Title : Senior DevOps Engineer

Department : Engineering & Infrastructure

Experience Required : 5- 8 years

Key Responsibilities :

1. Infrastructure Management (AWS) :


- Own and manage the entire AWS infrastructure - EC2, RDS, S3, ECS/EKS, VPC, IAM, CloudFront, Route 53, and related services.

- Design for high availability and fault tolerance; ensure infra can handle payment-grade uptime requirements.

- Right-size and optimise infrastructure for cost without compromising reliability.

- Maintain environment parity across dev, staging, and production.

2. CI/CD Pipelines :

- Build, maintain, and continuously improve CI/CD pipelines across all services supporting a daily release cadence.

- Manage deployment triggering across multiple AWS availability zones; handle traffic routing between zones including manual intervention when required.

- Ensure fast, reliable, and safe deployments with rollback capabilities and pre/post-release health checks.

- Work closely with the engineering team to reduce deployment friction and maintain release velocity.

- Manage branching strategies, environment promotion, and deployment gates.

3. Monitoring, Alerting & Incident Response :

- Own site monitoring end-to-end - set up and maintain CloudWatch and Site24x7 (or equivalent) for uptime checks, dashboards, log aggregation, and alerting.

- Configure intelligent alerting rules that surface real issues without creating noise; maintain runbooks for common alert scenarios.

- Be the first responder for production alerts - acknowledge, triage, and resolve or escalate within defined SLAs. Proactive action on alerts is a core expectation of this role.

- Conduct root cause analysis (RCA) for all production incidents and drive fixes to prevent recurrence.

- Maintain an on-call schedule; this role requires availability outside business hours for critical alerts.

4. System Automation :

- Automate repetitive operational tasks - provisioning, scaling, patching, backups, and configuration management.

- Manage infrastructure-as-code using Terraform or equivalent; no manual console changes in production.

- Automate monitoring setup, alerting thresholds, and runbook execution wherever possible.

5. Security & Compliance :

- Implement and maintain security rules, firewall policies, network ACLs, and IAM roles with least-privilege principles.

- Manage VPN setup and access controls for internal teams and vendors.

- Manage access provisioning and deprovisioning for all team members across AWS and related services.

- Ensure infrastructure compliance with PCI DSS and ISO 27001 requirements - work closely with the InfoSec Lead on audit evidence and remediation.

- Manage SSL/TLS certificates, key rotation, secrets management (AWS Secrets Manager / Vault), and encryption at rest and in transit.

6. Releases & Server Operations :

- Coordinate and execute daily production releases - validate pre-release checklists, trigger deployments across zones, manage traffic routing, and monitor post-release health.

- Handle server maintenance, OS upgrades, dependency patching, and scheduled downtime windows.

- Manage database backups, restoration drills, and disaster recovery procedures.

- Maintain up-to-date infrastructure documentation and runbooks.

7. AI-Augmented Development Infrastructure :

- Design and manage infrastructure for AI-assisted development workflows - including on-demand provisioning and teardown of ephemeral EC2 or container instances used by AI dev tools such as Claude Code.

- Integrate AI dev tooling into CI/CD pipelines - enabling automated code generation, review, and testing stages that spin up isolated compute, execute tasks, and clean up on completion.

- Implement IAM policies, network boundaries, and cost guardrails for ephemeral AI development instances.

- Build monitoring and observability into AI-augmented pipelines - tracking instance lifecycle, run times, failure rates, and compute costs.

Requirements :

Essential Experience :

- 5- 8 years of hands-on DevOps or infrastructure engineering experience.

- Deep, practical AWS expertise - EC2, RDS, ECS/EKS, VPC, IAM, ALB, Route 53, CloudWatch. Ability to architect, troubleshoot, and optimise independently.

- Proven experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, or equivalent) including multi-zone deployments.

- Hands-on experience with monitoring stacks - CloudWatch, Site24x7, Grafana, Prometheus, Datadog, or equivalent.

- Infrastructure-as-code using Terraform, CloudFormation, or Pulumi.

- Strong scripting skills - Python, Bash, or equivalent.

- Container experience - Docker and Kubernetes (EKS or self-managed).

- Experience managing VPN solutions (OpenVPN, AWS Client VPN, WireGuard, or equivalent).

- Access management experience - provisioning, deprovisioning, and periodic reviews across cloud and tooling.

Nice to Have :

- Experience integrating AI dev tools (Claude Code, GitHub Copilot, Cursor, or similar) into CI/CD pipelines.

- Familiarity with ephemeral environment patterns for automated testing or AI-assisted workflows.

- AWS certifications - Solutions Architect, DevOps Engineer Professional, or Security Specialty.

- Familiarity with PCI DSS infrastructure requirements and ISO 27001 technical controls.

- Experience in a fintech, payments, or BFSI environment.

The Ideal Candidate Profile :

What we really mean by 'reliable and dependable'

In a payments business, infrastructure failures have real, immediate consequences for merchants and customers. We need someone who treats a production alert at 11 PM the same way they would at 11 AM - someone who does not need to be chased, who over-communicates during incidents, and who takes pride in a green dashboard. If that describes you, we want to talk.

- You have a personal standard for uptime that is higher than what any SLA requires.

- You automate things the first time you have to do them manually twice.

- You document as you go - not as an afterthought.

- You are the person your engineering team trusts when things break at odd hours.

- A daily release cadence energizes you rather than stresses you - you have done this before, and you own it end-to-end.

- You are excited about building AI-augmented dev pipelines - integrating tools like Claude Code into CI/CD is exactly the kind of problem you enjoy solving.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...