HamburgerMenu
hirist

Senior DevOps & SecOps Engineer - Google Cloud Platform

Vedantu Innovations
4 - 6 Years
Bangalore

Posted on: 30/09/2026

Job Description

Role : Senior DevOps & SecOps Engineer

Location : Bengaluru


Experience : 4 - 6 years | Individual Contributor

About the Role :

Vedantu operates a mature, large-scale cloud infrastructure primarily on Google Cloud Platform (GCP), supporting our learning platforms, internal systems, data workloads, and engineering teams. Our infrastructure has evolved over several years and is currently in a stable state with established architecture, operational practices, monitoring, and deployment processes. Engineering leads and service owners actively own and understand the infrastructure associated with their applications.

We are looking for a Senior DevOps & SecOps Engineer who can take end-to-end operational ownership of this environment and ensure that our infrastructure remains reliable, secure, cost-efficient, well-governed, and up to date.

This is not a role where you will be expected to build an entire DevOps ecosystem from scratch or independently support every application. You will work closely with experienced engineering leads and service owners who understand their systems and infrastructure well. In addition to DevOps ownership, you will own infrastructure-focused security operations such as access governance, vulnerability management, hardening, security monitoring, patching, and audit support.

What You Will Own :

- Google Cloud Platform infrastructure

- Google Kubernetes Engine (GKE)

- Compute, networking, storage, CDN and load-balancing infrastructure

- IAM, service accounts and secrets management

- Kafka clusters

- GitLab and GitLab Runners, Argo CD

- Elasticsearch/Kibana stack

- Grafana and monitoring infrastructure

- CI/CD infrastructure

- Infrastructure automation

- Cloud security and access management

- Infrastructure vulnerability management and remediation

- Privileged access reviews and IAM governance

- Security monitoring and infrastructure-level incident response

- Security hardening, patching and audit evidence support

- Infrastructure upgrades and patching

- Production infrastructure troubleshooting

- Cloud usage and cost monitoring

- AWS infrastructure used by selected systems

Key Responsibilities :

Cloud Infrastructure :

Own and maintain our production infrastructure on Google Cloud Platform, including :

- Compute Engine

- Google Kubernetes Engine (GKE)

- Cloud Storage

- Load Balancing

- Cloud CDN

- VPC and networking

- Cloud NAT

- IAM and Service Accounts

- Secret Manager

- DNS and related networking services

- Data and infrastructure services used by engineering teams

You should be comfortable understanding an existing cloud architecture, troubleshooting it, making changes safely, and improving it where required.

Kubernetes & Containers :

- Operate and maintain production Kubernetes workloads on GKE.

- Troubleshoot Kubernetes, networking, resource utilization, scaling and deployment issues.

- Manage cluster upgrades and infrastructure changes.

- Work with engineering teams on containerization and deployment-related issues.

- Ensure appropriate resource allocation, availability and reliability of workloads.

CI/CD & Developer Infrastructure :

Maintain and improve our engineering infrastructure, including GitLab, GitLab Runners, Argo CD, CI/CD pipelines, and build and deployment infrastructure.

Work with engineering teams to troubleshoot deployment failures and improve the reliability and efficiency of our CI/CD systems.

Monitoring, Logging & Observability :

- Own and maintain Grafana, Kibana / Elasticsearch, application and infrastructure monitoring, and alerting.

- Ensure production systems have appropriate monitoring and infrastructure issues can be detected and diagnosed quickly.

Kafka & Self-Hosted Infrastructure :

Maintain self-managed infrastructure such as our open-source Kafka clusters and other internally hosted engineering systems.

- Version upgrades

- Capacity and resource monitoring

- Troubleshooting

- Availability and performance

- Configuration management

- Backup and recovery practices where applicable

Security Operations (SecOps) :

Own infrastructure and cloud security operations as an extension of the DevOps role. This includes :

- Review and maintain GCP IAM, service accounts, privileged roles and least-privilege access.

- Conduct periodic access reviews across GCP, GKE/Kubernetes, GitLab, databases and infrastructure tooling.

- Monitor infrastructure vulnerabilities and coordinate timely remediation of critical and high-risk findings.

- Manage OS, container image and infrastructure security patching.

- Review firewall rules, public exposure, network access and cloud configurations for unnecessary security risk.

- Maintain Kubernetes security controls including RBAC, workload permissions, secrets and cluster hardening.

- Own certificate, key and secrets hygiene, including expiry monitoring and rotation of stale credentials.

- Monitor cloud audit logs and infrastructure security alerts and investigate suspicious activity.

- Support infrastructure-level security incidents, root-cause analysis and remediation.

- Periodically harden self-hosted systems such as GitLab, Kafka, Elasticsearch/Kibana, Grafana and VMs.

- Provide infrastructure evidence and remediation support for security and compliance audits.

- Automate recurring security checks, configuration validation, vulnerability detection and patching wherever practical.

The role is focused on cloud and infrastructure security operations. Application security, penetration testing, privacy/legal compliance and broader security governance are not primary ownership areas for this position.

Infrastructure Maintenance & Upgrades :

- Kubernetes upgrades

- GitLab upgrades

- Grafana/Kibana/Elasticsearch upgrades

- OS and security patching

- Dependency and infrastructure upgrades

- Certificate and secret management

- Identifying obsolete infrastructure

- Capacity management

- Upgrades should be planned and executed with minimal production impact.

Production Reliability :

Participate in troubleshooting infrastructure-related production incidents.

You should be comfortable diagnosing problems across Application - Kubernetes - Compute - Network - Load Balancer - Storage - Database / Messaging - Cloud Infrastructure.

Work closely with application owners and engineering leads during incidents and help identify root causes and preventive actions.

Cloud Cost Management :

- Understand major infrastructure cost drivers.

- Identify unusual increases in cloud spend.

- Map infrastructure usage to services/workloads.

- Identify idle or over-provisioned resources.

- Work with engineering teams on cost optimization.

- Support cloud invoice and usage reconciliation.

The objective is to maintain the right balance between reliability, performance and cost.

What We Are Looking For :

Must Have :

- 4 - 6 years of hands-on DevOps / Cloud Infrastructure / SRE experience.

- Strong production experience with Google Cloud Platform (GCP).

- Strong hands-on experience with Kubernetes, preferably GKE.

- Strong Linux administration and troubleshooting skills.

- Good understanding of cloud networking including VPCs, routing, NAT, DNS, load balancers and firewalls.

- Experience operating production systems with high availability requirements.

- Strong understanding of IAM, service accounts, secrets and cloud security.

- Experience with Docker and containerized applications.

- Experience managing CI/CD infrastructure.

- Experience with monitoring and observability platforms such as Grafana, Prometheus, Elasticsearch and Kibana.

- Experience with Infrastructure as Code, preferably Terraform.

- Strong scripting skills using Python, Shell/Bash or similar languages.

- Ability to independently troubleshoot production infrastructure issues.

- Hands-on experience with cloud security, IAM governance, vulnerability management and infrastructure hardening.

- Understanding of Kubernetes security, RBAC, secrets management and container security.

- Experience handling infrastructure security alerts, remediation and access reviews.

Good to Have :

- Experience managing Kafka in production, particularly self-hosted Kafka.

- Experience administering GitLab and GitLab Runners, Argo CD

- Experience operating Elasticsearch/Kibana.

- Experience with AWS services such as EC2, S3, IAM, SQS, SNS and VPC.

- Experience with GCP data infrastructure and managed data-processing services.

- Experience with databases such as MongoDB and MySQL.

- Experience with infrastructure security, vulnerability management and security patching.

- GCP or Kubernetes certifications.

- Experience with GCP Security Command Center or equivalent cloud security tooling.

- Experience with vulnerability scanners, container/image scanning and security automation.

- Exposure to security/compliance audits and evidence collection.

What Makes Someone Successful in This Role :

- Can independently take ownership of a mature production infrastructure.

- Is comfortable inheriting and understanding existing systems rather than wanting to rebuild everything.

- Knows when infrastructure actually needs improvement and when a stable system should simply be maintained well.

- Is systematic about upgrades, patching, monitoring and operational hygiene.

- Can troubleshoot deeply when something goes wrong.

- Works collaboratively with engineering leads and service owners.

- Automates repetitive operational work instead of repeatedly performing it manually.

- Thinks about infrastructure reliability, security and cost together.

- Treats security hygiene, access reviews, patching and vulnerability remediation as ongoing operational responsibilities.

- Documents changes and keeps infrastructure knowledge distributed across the engineering team.

Our Environment :

Our primary cloud is GCP, where we use a broad range of compute, networking, storage, Kubernetes, security and data services. Our engineering infrastructure also includes :

GCP | GKE | Kubernetes | Docker | Kafka | GitLab | GitLab Runners | Argo CD| Grafana | Prometheus | Elasticsearch | Kibana | Terraform | Linux | MongoDB | MySQL | AWS SQS/SNS | IAM | Secrets Management | Vulnerability Management

You will inherit a stable environment with established engineering ownership and experienced technical leads who understand their respective systems. Your responsibility will be to keep this infrastructure reliable, secure, current and efficient while helping us evolve it as our engineering needs change.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...