HamburgerMenu
hirist

Techgrit - Migration Specialist - Observability & Automation

TechGrit
3 - 8 Years
Remote

Posted on: 22/09/2026

Job Description

Role : Migrations Specialist

Location : Remote

Experience : 2 Years +

Department : Customer Experience / Site Reliability Engineering

Role Overview :

We are hiring a Migrations Specialist for our client, a leader in software quality and security management offering a comprehensive suite of testing, security, and product delivery tools. This is a hands-on, execution-focused role supporting established cloud infrastructure and SaaS platforms across the client's production and staging environments. You will operate within defined architectural patterns and reliability frameworks while helping improve observability, automation, and operational maturity. The ideal candidate is comfortable responding to live production events, managing operational requests, and continuously improving systems through automation.

Key Responsibilities :

Production & Incident Management :

- Monitor and support production and staging environments in real time, ensuring high availability, performance, and stability.

- Respond to incidents, perform triage and root cause analysis, and contribute to post-incident reviews and remediation efforts.

- Participate in an on-call rotation with defined SLAs.

- Handle ad-hoc and unplanned operational requests from Product, Support, and internal teams.

Observability & Automation :

- Maintain and enhance monitoring, alerting, dashboards, logs, and metrics; improve signal-to-noise ratio and standardize observability practices.

- Contribute to automation efforts to reduce operational toil.

Delivery & Infrastructure :

- Support CI/CD pipelines, production releases, and GitOps workflows.

- Maintain and improve Kubernetes-based infrastructure and containerized workloads.

- Support Infrastructure as Code practices and ongoing environment improvements.

Migrations :

- Migrate applications between environments; these migrations typically take place on weekends, so weekend availability is required.

Detailed Requirements :

- 2+ years of experience in Site Reliability Engineering, DevOps, or Production Operations.

- Demonstrable AWS experience supporting production environments.

- Experience supporting production SaaS applications.

- Strong understanding of CI/CD systems (GitHub Actions, Jenkins, CircleCI, or similar).

- GitOps experience and strong Git fundamentals.

- Experience using GitHub, Jira, and Confluence in collaborative engineering environments.

- Kubernetes experience (EKS, kOps, or similar).

- Docker/containerization experience.

- Observability stack experience (Grafana, Prometheus, Loki, PagerDuty, or similar).

- Scripting experience (Bash, Python, or Go).

- Working knowledge of relational databases (e.g., PostgreSQL, MySQL) for troubleshooting and operational support.

- Infrastructure as Code experience (Terraform, Helm, or similar).

- Comfortable working within structured operational processes and SLAs.

- Strong written and verbal English communication skills; able to clearly explain technical concepts.

- Self-driven with a growth mindset

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...