Posted on: 08/06/2026
Job Description :
We are seeking a highly experienced Principal DevOps Engineer to lead the design, evolution, and scalability of our cloud infrastructure and platform engineering initiatives. This role requires a technical leader who has successfully managed large-scale, multi-tenant SaaS platforms and driven infrastructure strategy across multiple engineering teams.
The ideal candidate will possess deep expertise in AWS, Kubernetes, Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Automation, and Cloud-Native Technologies. As a Principal Engineer, you will define the long-term platform roadmap, establish engineering standards, and ensure reliability, scalability, security, and operational excellence across the organization.
Key Responsibilities :
- Define and execute the organization's cloud infrastructure and platform engineering strategy.
- Architect and manage highly available, scalable, and secure multi-tenant SaaS platforms.
- Drive platform modernization initiatives and cloud-native transformation programs.
- Establish infrastructure standards, governance frameworks, and operational best practices.
- Collaborate with Engineering, Product, Security, and Architecture teams to align platform capabilities with business objectives.
- Design and optimize AWS-based infrastructure leveraging services such as :
1. EKS (Elastic Kubernetes Service)
2. EC2
3. RDS
4. VPC
5. Route53
6. ELB/ALB
7. IAM
8. CloudWatch
9. S3
10. Lambda
- Build resilient and cost-efficient cloud environments capable of supporting rapid business growth.
- Lead infrastructure capacity planning, disaster recovery, and business continuity initiatives.
- Architect and operate Kubernetes clusters at scale across production and non-production environments.
1. Implement container orchestration strategies for high availability and performance.
2. Establish standards for cluster management, service mesh, ingress controllers, and workload optimization.
3. Improve deployment reliability, scalability, and security of containerized applications.
- Define Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs).
1. Lead reliability engineering initiatives to achieve platform uptime and performance targets.
2. Build proactive monitoring, alerting, observability, and incident management frameworks.
3. Conduct post-incident reviews and drive continuous improvement initiatives.
- Develop and optimize CI/CD pipelines for large-scale SaaS environments.
1. Drive Infrastructure as Code (IaC) adoption using tools such as Terraform, CloudFormation, or Pulumi.
2. Automate provisioning, deployments, configuration management, and operational workflows.
3. Eliminate manual operational overhead through self-service platform capabilities.
- Embed security best practices throughout infrastructure and deployment processes.
1. Implement cloud security controls, secrets management, access governance, and vulnerability management.
2. Support compliance initiatives including SOC2, ISO 27001, GDPR, HIPAA, or similar frameworks.
3. Partner with security teams to strengthen platform security posture.
- Act as the highest level technical authority for DevOps and Platform Engineering.
1. Mentor Senior Engineers, Staff Engineers, and Engineering Managers.
2. Drive engineering excellence, architectural reviews, and technical decision-making.
3. Influence platform roadmap and strategic technology investments across the organization.
Required Skills & Qualifications :
- 10-15 years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or related domains.
- Experience working in B2B SaaS Product Companies is mandatory.
- Proven experience in roles such as :
1. Principal DevOps Engineer
2. Staff DevOps Engineer
3. Lead DevOps Engineer
4. Principal SRE
5. Staff SRE
6. Platform Engineering Lead
- Demonstrated ownership of organization-wide infrastructure strategy and platform roadmap.
- Expert-level experience with AWS Cloud services :
1. EKS
2. EC2
3. RDS
4. VPC
5. IAM
6. Route53
7. CloudWatch
8. S3
9. Networking Services
- Strong understanding of distributed systems architecture.
- Deep expertise in Kubernetes administration and architecture at scale.
- Strong hands-on experience with :
1. Docker
2. Kubernetes
3. Helm
4. Service Mesh technologies
5. Container Security
- DevOps & Automation
1. Extensive experience with :
i. Terraform
ii. CloudFormation
iii. Ansible
iv. GitOps methodologies
2. CI/CD tools :
i. Jenkins
ii. GitHub Actions
iii. GitLab CI/CD
iv. ArgoCD
v. CircleCI
- Monitoring & Observability
1. Expertise in :
i. Prometheus
ii. Grafana
iii. ELK Stack
iv. Datadog
v. New Relic
vi. OpenTelemetry
- Strong incident management and production support experience.
- Networking & Security
1. Advanced understanding of :
i. VPC Design
ii. Load Balancing
iii. DNS
iv. VPN
v. CDN
vi. Network Security
vii. Zero Trust Architecture
- Experience implementing cloud security best practices and compliance controls.
- Experience supporting large-scale multi-tenant SaaS applications serving enterprise customers.
- Experience building Internal Developer Platforms (IDP).
- Knowledge of FinOps and cloud cost optimization.
- Experience with multi-region deployments and global infrastructure.
- AWS Professional Certifications preferred.
- Exposure to AI/ML infrastructure workloads is a plus.
- Strategic Infrastructure Leadership
- Platform Engineering Excellence
- Cloud-Native Architecture
- Reliability Engineering
- Technical Mentorship
- Cross-Functional Stakeholder Management
- Problem Solving & Decision Making
- Scalability & Performance Optimization
- Security & Compliance Mindset
- 10-15 years of total experience.
- Current or recent experience in a B2B SaaS Product Company.
- Hands-on expertise with AWS, Kubernetes, DevOps, SRE, and Platform Engineering.
- Experience operating multi-tenant SaaS platforms at scale.
- Proven track record of leading infrastructure strategy and platform roadmap at an organizational level.
- Experience in Principal, Staff, or Lead-level engineering roles.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1642573