Senior DevOps Engineer - AWS & Kubernetes
Job Summary:
We are looking for an experienced Senior DevOps Engineer to own the reliability, scalability, security, and automation of production platforms running on AWS and Kubernetes.
The ideal candidate will have strong hands-on experience in AWS, Kubernetes, Terraform, Docker, Jenkins, GitOps, Linux, Prometheus, and Grafana, along with proven experience supporting production environments and participating in incident response/on-call rotations.
This is a hands-on engineering role with significant ownership of cloud infrastructure, CI/CD, Infrastructure as Code, observability, security, and platform reliability.
Key Responsibilities:
Cloud Infrastructure & Platform Operations:
- Own the reliability and availability of production workloads running on AWS and Kubernetes.
- Operate and troubleshoot Kubernetes clusters in production environments.
- Manage and maintain Linux servers, including Amazon Linux and Rocky Linux.
- Perform incident triage, troubleshooting, root-cause analysis, and remediation.
- Monitor cloud capacity, performance, and cost optimization.
- Work closely with engineering teams to improve platform scalability and reliability.
CI/CD & Infrastructure as Code :
- Design, maintain, and improve end-to-end CI/CD pipelines using Jenkins.
- Drive the transition toward GitOps-based deployment and automation.
- Work with FluxCD and other GitOps technologies.
- Develop and maintain Infrastructure as Code using Terraform.
- Manage Kubernetes configurations using Kustomize.
- Improve developer experience by enabling reliable and automated deployment processes.
Monitoring & Observability :
- Maintain and enhance monitoring and observability using Prometheus and Grafana.
- Build dashboards, alerts, metrics, and SLOs.
- Identify performance and reliability issues proactively.
- Establish effective monitoring and alerting for production services.
Incident Management & Reliability :
- Participate in a scheduled on-call rotation.
- Lead production incident response and service restoration.
- Perform RCA (Root Cause Analysis) and drive corrective actions.
- Conduct blameless post-mortems and ensure follow-up actions are completed.
- Develop and maintain operational runbooks and documentation.
Security & Compliance :
- Implement secure-by-default practices across cloud infrastructure, CI/CD pipelines, containers, and Kubernetes.
- Support security reviews, access management, and secrets management.
- Contribute to compliance requirements including ISO 27001, PCI-DSS, and Cyber Essentials.
Engineering & Continuous Improvement :
- Mentor engineers on DevOps, infrastructure, deployment, and reliability practices.
- Contribute to platform architecture and technical decisions.
- Evaluate and introduce tools and automation that provide measurable value.
- Maintain clear technical documentation, runbooks, and operational procedures.
Required Skills & Experience :
- Strong hands-on experience managing AWS production environments.
- Proven experience operating Kubernetes in production, preferably Amazon EKS.
- Strong knowledge of Docker and containerized applications.
- Hands-on experience with Terraform / Infrastructure as Code.
- Strong experience with Jenkins and CI/CD pipelines.
- Experience with GitOps, preferably FluxCD or Argo CD.
- Good understanding of Linux administration, preferably Amazon Linux/Rocky Linux.
- Strong understanding of HTTP, networking, DNS, TCP/IP, and cloud infrastructure.
- Good scripting skills in Python and/or Bash/Shell.
- Hands-on experience with Prometheus and Grafana.
- Experience with production troubleshooting, incident management, RCA, and on-call support.
Technical Skills :
- Cloud: AWS, GCP, DigitalOcean
- Containers: Kubernetes, EKS, Docker
- IaC: Terraform, Kustomize
- CI/CD: Jenkins, GitOps, FluxCD, Argo CD
- Monitoring: Prometheus, Grafana
- OS: Linux, Amazon Linux, Rocky Linux
- Scripting: Python, Bash, Shell
- Database: MySQL
- Security: ISO 27001, PCI-DSS, Cyber Essentials
- Other: Helm, DevSecOps, SRE, Incident Management
Key Attributes :
- Strong ownership and problem-solving mindset.
- Reliability-focused and comfortable working with production systems.
- Pragmatic approach to engineering and automation.
- Strong troubleshooting and analytical skills.
- Clear written and verbal communication.
- Comfortable working independently and collaborating with cross-functional engineering teams.
- Willingness to participate in a paid weekly on-call rotation.
Experience :
- 6+ years of overall IT experience, with strong hands-on experience in DevOps, Cloud, SRE, Platform Engineering, or Infrastructure.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1669811