HamburgerMenu
hirist

Job Description

We are looking for a Senior Cloud Operations Engineer to support and optimize a highly available, scalable, and secure cloud infrastructure.

The ideal candidate should have strong expertise in cloud platforms, Linux administration, automation, container technologies, and production operations, with experience managing enterprise-scale cloud environments.

Key Responsibilities :

- Monitor, maintain, and support production cloud environments to ensure high availability, performance, and reliability.

- Manage and troubleshoot cloud infrastructure, Kubernetes clusters, Linux servers, Azure SQL, Kafka, MongoDB, and related services.

- Respond to production incidents, perform root cause analysis (RCA), and drive timely resolution.

- Automate infrastructure provisioning, deployments, and operational processes using Infrastructure as Code (IaC) and scripting.

- Collaborate with DevOps, Engineering, and Support teams to deploy applications and improve operational efficiency.

- Monitor system health, capacity, performance, and availability, ensuring operational SLAs are consistently met.

- Build and enhance operational processes to improve platform stability, scalability, and security.

- Implement infrastructure upgrades, patching, and lifecycle management following industry best practices.

- Create and maintain operational documentation, runbooks, and monitoring dashboards.

- Participate in on-call support and contribute to continuous service improvement initiatives.

Required Skills :

- 5 - 10 years of experience in Cloud Operations, Site Reliability Engineering (SRE), DevOps, or Infrastructure Operations.

- Hands-on experience with Microsoft Azure (preferred) or AWS/GCP.

- Strong Linux system administration and troubleshooting skills.

- Experience managing Kubernetes and containerized applications.

- Experience with Azure SQL, Kafka, MongoDB, and distributed systems.

- Hands-on experience with Infrastructure as Code tools such as Terraform.

- Experience with configuration management and automation tools such as Ansible or Chef.

- Strong scripting skills in Bash, PowerShell, or Python.

- Good understanding of CI/CD pipelines and DevOps practices.

- Experience with monitoring, alerting, performance tuning, and incident management.

- Strong analytical, troubleshooting, communication, and stakeholder management skills.

Good to Have :

- Experience with cloud-native architectures and microservices.

- Exposure to observability tools such as Prometheus, Grafana, Datadog, or Azure Monitor.

- Knowledge of networking, security, and disaster recovery best practices.

- Experience working in Agile and ITIL-based operational environments.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...