Responsibilities :
Infrastructure at Scale :
- Design and evolve our cloud-native infrastructure (AWS/Kubernetes), ensuring availability, performance, and cost efficiency across regions and products.
Platform & Developer Experience :
- Build internal tools and platforms that help engineers deploy, monitor, and scale their services independently with minimal friction and maximum confidence.
CI/CD & Release Automation :
Architect secure, fast, and scalable CI/CD pipelines across multiple environments using tools like GitHub Actions, and Jenkins.
Reliability Engineering :
Champion observability, SLOs, and incident response practices. Drive a culture of proactive performance monitoring and resilient system design.
Security & Governance :
Integrate DevSecOps practices from policy-as-code and automated audits to secure secrets management and vulnerability scanning.
Mentorship & Thought Leadership :
Guide and mentor DevOps and SRE engineers. Partner closely with platform developers on infrastructure strategy, deployment patterns, and production readiness.
Ideal Candidate :
Strong Principal DevOps Engineer Profile :
- Must have 10+ years in DevOps / SRE / Infrastructure roles with hands-on experience (clear scale signals like traffic, uptime, latency, infra size should be mentioned) in B2B SAAS companies
- Must have worked in Principal / Staff / Lead DevOps / SRE / Platform Engineer role and demonstrated org-level ownership - setting infra roadmap, defining DevOps charter, or structuring the platform function not just domain-level technical ownership
- Must show evidence of strategic authorship, defined multi-year infra/platform strategy, drove company-wide architectural shifts as an initiator (not implementer), or directly interfaced with VP Eng / CTO / product leadership on infra direction
- Must have B2B SaaS company experience with multi-tenant architecture OR multiple production stacks (multi-env / multi-client systems)
- AWS (VPC, EKS, EC2, RDS, networking), Kubernetes (EKS) at scale, Designing high availability, multi-region systems
- Automation & IaC) : Terraform (must-have), Helm / GitOps, Strong scripting (Python / Go / Bash)
- Mandatory (Tech Skills 4 - Reliability & Observability) : SRE principles (SLOs, SLIs, error budgets), Monitoring tools (Prometheus, Grafana, Datadog), Alerting, on-call, incident management
- Must demonstrate leadership experience in an individual contributor capacity having mentored senior engineers, driven cross-team technical alignment, or anchored org-wide initiatives without having moved into a people management or engineering manager role
- Strong B2B SaaS product companies only
Preferred (Education) :
- B.Tech in Computer Science or related fields