HamburgerMenu
hirist

Job Description

Role : DevOps & Automation Engineer (PagerDuty Integration) :

Openings : Lead / Principal DevOps Engineer (10+ years)

Location : Delhi NCR

Work model : Hybrid, Gurgaon two days a week

Shift : Rotational

Experience : 10+ years

Must have skills :

- Lead / Principal DevOps Engineer (10+ years) with Lead exp

- Candidate must have 3 yrs Lead exp

- Python scripting with good automation

- Need hands on experience in any two clouds (AWS / GCP / Azure)

- Excellent communication is required

Key Responsibilities :

- Design, implement, and maintain PagerDuty / incident.io / Rootly / opsgenie (Atlassian) / squadcast, integrations with monitoring and observability tools (e.g., Amazon CloudWatch, Datadog, Prometheus/Alertmanager, Grafana, New Relic, Splunk, Dynatrace, Nagios/Zabbix).

- Automate, onboarding, configuration management, and incident response workflows using Terraform, PagerDuty APIs, and Rundeck-based operational automation.

- Configure PagerDuty / incident.io / Rootly / opsgenie (Atlassian) / squadcast services, escalation policies, schedules, event orchestration, routing rules, and alert grouping and suppression to reduce alert fatigue.

- Build event-driven automation on AWS/Azure/GCP (EventBridge, SNS, Lambda, Systems Manager, Step Functions) for auto-remediation and incident enrichment.

- Integrate PagerDuty / incident.io / Rootly / opsgenie (Atlassian) / squadcast with ITSM and collaboration tools such as ServiceNow, Jira, Slack, monitoring, cloud, identity and data platforms and Microsoft Teams for ticketing and incident communication.

- Build and maintain CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI, or AWS CodePipeline) to deliver monitoring and alerting configuration as code.

- Create runbooks, integration documentation, and standard operating procedures.

- Collaborate with the client's SRE, application, platform, and security teams to onboard services onto standardised alerting and on-call practices.

Required Skills :

- Hands-on experience with PagerDuty / incident.io / Rootly / opsgenie (Atlassian) / squadcast (services, integrations, escalation policies, event orchestration, Events API v2, REST API).

- Strong Cloud infrastructure and DevOps experience.

- Proficiency in Infrastructure as Code, preferably Terraform.

- Requirement for adaptability across platforms and tools, not restricted to a single stack.

- Strong automation and scripting skills in Python, Bash, JavaScript (for webhooks, event transformers, and payload handling), including REST API and webhook integrations.

- Should have strong understanding of AIOPs concepts.

- Experience with monitoring and observability tools and with building meaningful alerts and SLO-based alerting.

- Working knowledge of CI/CD tooling and Git-based workflows.

- Solid understanding of incident management, on-call practices, and SRE principles.

The job is for:

May work from home
info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...