Posted on: 08/07/2026
Job Description :
We are looking for a hands-on CloudOps engineer with strong experience supporting production environments across AWS and GCP.
Role & Responsibilities :
Cloud Platforms (AWS & GCP) :
Must have hands-on operational experience with :
- AWS (strong) : EC2, ECS/Fargate
- AWS (working knowledge) : RDS (Postgres/MySQL)
- GCP (strong) : GCE (VMs)
- GCP (working knowledge) : CloudSQL, AlloyDB
Expected Capabilities :
- Troubleshoot performance, availability, and connectivity issues.
- Support upgrades, patching coordination, and lifecycle management.
- Work across dev, QA, and production environments with strong awareness of production impact.
Networking (Practical Knowledge) :
Strong working understanding of :
- VPCs, subnets, routing tables
- Security groups and firewall rules
- DNS (internal and external)
- Basic connectivity patterns (peering, private endpoints / PSC)
Ability to :
- Troubleshoot cross-service connectivity issues.
- Collaborate effectively with networking teams (clear understanding of ownership boundaries).
IAM, Security & Access Governance (Critical) :
Required experience with :
- AWS IAM roles and policies
- GCP IAM (projects, folders, service accounts)
Must understand :
- Least privilege access model
- PAM / temporary elevation concepts
- Service account usage and restrictions
Infrastructure as Code (IaC) :
Required :
- Terraform (hands-on)
Strong plus :
- AWS CloudFormation
Expectation :
- Comfortable working in environments where all changes go through pipelines (no manual console changes).
Monitoring, Logging & Observability :
Experience with :
- AWS CloudWatch
- GCP Monitoring and Logging
- Splunk or similar tools
Ability to :
- Troubleshoot incidents using logs and metrics.
- Tune alerts and reduce noise.
DevOps Tooling Support :
Must have experience supporting :
- Artifactory
- Terraform pipelines
Familiarity with :
- GitHub / repository management
- Docker / Kubernetes (basic understanding)
Incident Management & Operations :
Required experience with :
- On-call support model
- Incident triage, escalation, and ownership
Expectations :
- Own issues end-to-end (triage - resolution - handoff if needed).
- Communicate clearly during incidents.
- Work across Cloud Engineering, application teams, and vendors.
Documentation & Runbooks :
Strong expectation for :
- Writing and maintaining runbooks.
- Keeping documentation accurate and up to date.
Ability to :
- Convert informal knowledge into structured, usable documentation.
Automation & Continuous Improvement :
Experience with :
- Scripting (Python or Bash).
- Automating operational tasks and workflows.
Mindset :
- Focus on reducing manual effort and improving operational efficiency.
Cloud Lifecycle Management :
Experience with :
- Provisioning, upgrades, and decommissioning.
- OS / image management (e.g., AMIs, VM images).
Understanding of :
- Patching responsibilities (shared vs application-owned).
- Version upgrades (e.g., database upgrades).
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1652349