HamburgerMenu
hirist

Job Description

We're looking for an experienced SRE / Incident Response Engineer to join our team . You'll take point on produc%on incident response, deep-dive troubleshoo-ng, and light automa-on work, suppor-ng a team of junior SRE/Batch Ops analysts.

If you thrive in high-urgency environments, know your way around modern observability stacks, and can bring structure + calm to incidents, we want to talk to you.

What You'll Do :

- Lead real- me incident response, triage, escala-on, and stakeholder comms Run post- incident reviews and drive follow-through on ac-on items


- Monitor and troubleshoot systems using Datadog (logs, APM, dashboards) Support batch workflows using JAMS / Control-M or similar schedulers Define & measure SLIs/SLOs with engineering teams


- Build automation in Python or PowerShell to reduce operational toil

- Partner with plaUorm & product teams to improve reliability and integrations Participate in the 24/7 on-call rotation (shared with junior staff)


Skills & Experience Must-Have :

- 2 to 4 years in SRE, production ops, or incident response

- Strong incident command + cross-team communication skills Datadog proficiency (APM, metrics, log analysis)

- Familiarity with JAMS, Control-M, or enterprise batch schedulers Basic AWS, SQL, GitHub workflows

- Python or PowerShell scripting

- ITIL fundamentals (incident/problem/change) Containers/Kubernetes familiarity

Nice to Have :

- AIOps experience (Datadog AI, BigPanda, Moogsoc, PD AIOps, etc.) Exposure to LLM- based ops tools, runbook automa-on, or copilots PlaUorm engineering or DevOps tooling experience

- AWS or Kubernetes certifications

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...