HamburgerMenu
hirist

Job Description

Description :

Role & Responsibilities :


- Manage and maintain highly available, scalable, and reliable production infrastructure.

- Monitor system performance, uptime, latency, and application health across environments.

- Build and improve CI/CD pipelines for faster and safer deployments.

- Automate infrastructure provisioning, deployments, monitoring, and operational tasks.

- Troubleshoot production incidents, outages, and performance bottlenecks.

- Lead incident response, root cause analysis (RCA), and postmortem activities.

- Implement Infrastructure as Code (IaC) using tools like Terraform or Ansible.

- Manage Kubernetes clusters, container orchestration, and cloud-native platforms.

- Define and maintain SLI/SLO/SLA standards and reliability metrics.

- Optimize cloud infrastructure for scalability, performance, and cost efficiency.

- Collaborate with development, DevOps, security, and platform teams to improve system reliability.

- Ensure backup, disaster recovery, failover, and business continuity strategies are in place.

- Improve observability using monitoring, logging, and alerting tools.

- Participate in on-call rotations and production support activities.

- Mentor junior engineers and drive operational best practices.

Preferred Candidate Profile :


- Bachelors degree in Computer Science, Information Technology, or related field.

- 5 to 10 years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, or Platform Engineering.

- Strong hands-on experience with AWS, GCP, or Azure cloud platforms.

- Expertise in Kubernetes, Docker, and containerized environments.

- Strong experience with Terraform, Ansible, or other Infrastructure as Code tools.

- Proficiency in Linux administration and shell scripting.

- Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or ArgoCD.

- Knowledge of monitoring and observability tools like Prometheus, Grafana, Datadog, ELK, or Splunk.

- Good programming/scripting skills in Python, Go, or Bash.

- Strong understanding of networking concepts, security practices, and distributed systems.

- Experience handling production incidents and high-availability systems.

- Familiarity with microservices architecture and cloud-native technologies.

- Strong analytical, troubleshooting, and problem-solving abilities.

- Excellent communication and collaboration skills.

- Experience in 24x7 production support and on-call environments preferred.

Nice-to-Have Skills :


- Service Mesh (Istio/Linkerd)

- Kafka or distributed messaging systems

- FinOps / cloud cost optimization

- Chaos engineering

- MLOps or platform engineering exposure

- Multi-cloud infrastructure management

- DevSecOps practices

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...