- Ensure high system reliability, uptime, and performance through monitoring and alerting (Prometheus, Grafana, etc.)
- Perform infrastructure upgrades, patching, and cost optimization across environments
- Troubleshoot production issues and ensure quick resolution with minimal downtime
- Manage secrets, access controls, and cloud security best practices (IAM, networking, encryption)
- Collaborate closely with engineering and product teams to enable faster and safer releases
- Document infrastructure, processes, and best practices for scalability and maintainability
Ideal Candidate :
- Strong Devops/Cloud engineering profile
- Must have 6+ years of hands-on experience in DevOps or Cloud Engineering with at least 3+ years in a product-based or high-scaling startup environment
- Must have managed infrastructure at high scale or in fast-moving startup environments with full end-to-end ownership of one or more products
- Must have strong hands-on AWS experience as the primary cloud platform with exposure to multi-cloud environments
- Must have deep working knowledge of Kubernetes covering cluster management, deployments, and Helm charts for microservices
- Must have strong hands-on Terraform experience as the primary IaC tool with additional exposure to CloudFormation
- Must have experience building and managing CI/CD pipelines using GitHub Actions, GitLab CI, or equivalent tools
- Must have working experience with Prometheus, Grafana, ELK stack, or comparable monitoring and logging tooling
- Must have solid understanding of cloud networking fundamentals and cloud security best practices covering IAM, secrets management, and access controls
- Must be proficient in Bash and/or Python for automation, tooling, and infrastructure management
- Must have direct experience owning production environments including live troubleshooting, incident response, and minimal-downtime resolution