Posted on: 28/05/2026
Description :
You will not just react to incidents; you will actively engineer them out of existence by building resilient monitoring frameworks, automating repetitive operational tasks (to eliminate toil), and optimizing our cloud infrastructure and deployment pipelines.
Key Responsibilities :
- Drive end-to-end resolution of high-priority production issues, ensuring tight alignment with SLA/OLA targets.
- Lead post-mortem investigations following major incidents to identify structural bugs, code flaws, or infrastructure vulnerabilities, preventing recurrence.
- Ensure absolute compliance with standard ITIL/ITSM processes covering Incident, Problem, Change, and Release Management frameworks.
- Troubleshoot complex, multi-tiered enterprise application stacks hosted across Linux environments.
- Utilize intermediate-to-advanced Linux CLI commands and Shell Scripting alongside SQL queries to extract diagnostic data from databases and file systems.
- Support and optimize large-scale big data environments including Apache NiFi, Hadoop, and distributed data frameworks.
- Design and build custom alerts, transaction dashboards, and synthetic monitoring steps within Enterprise Observability tools (Splunk, Dynatrace, or similar APMs).
- Contribute to building centralized event-driven telemetry frameworks that correlate logs, traces, and metrics to trigger automated self-healing scripts.
- Write and maintain reusable, scalable infrastructure-as-code and configuration scripts using Ansible or Chef to standardize production environments.
- Manage, optimize, and secure automated Jenkins CI/CD pipelines using Groovy DSL and YAML workflow configurations.
- Standardize source code practices and branching strategies within distributed environments like Git /Bitbucket
- Monitor, scale, and optimize core cloud infrastructure components hosted on Microsoft Azure.
- Track cloud system capacity, identify performance bottlenecks, and recommend architecture optimizations to enhance system reliability.
Required Skills & Qualifications :
- Operating Systems & Scripting : Expert-level command of Linux systems coupled with hands-on Bash/Shell scripting for operations automation.
- Strong proficiency in Jenkins (Pipeline-as-code, Groovy, YAML) and Git / Bitbucket.
- Hands-on engineering experience using infrastructure automation tools like Ansible or Chef.
- Proven operational familiarity with Microsoft Azure (VMs, Networking, Storage, Monitoring).
- Robust experience navigating platform metrics and building analytical dashboards in Splunk or Dynatrace.
- Sound capability in writing SQL scripts for system verification and backend health audits.
- Clear structural understanding of ITIL/ITSM best practices.
- Deep background operating concurrently across Production Support and DevOps Delivery frameworks.
- Familiarity with Apache NiFi pipelines and Hadoop clustering.
- Calm under pressure; capable of leading high-stakes incident bridge calls with cross-functional technical teams.
- Intolerant of repetitive manual work; constantly seeking opportunities to automate operational tasks.
- Ability to clearly articulate complex infrastructure outages to senior stakeholders in plain, business-centric terms.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1639686