HamburgerMenu
hirist

Job Description

Job Description :

We are seeking a highly experienced Principal DevOps Engineer to lead the design, evolution, and scalability of our cloud infrastructure and platform engineering initiatives. This role requires a technical leader who has successfully managed large-scale, multi-tenant SaaS platforms and driven infrastructure strategy across multiple engineering teams.

The ideal candidate will possess deep expertise in AWS, Kubernetes, Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Automation, and Cloud-Native Technologies. As a Principal Engineer, you will define the long-term platform roadmap, establish engineering standards, and ensure reliability, scalability, security, and operational excellence across the organization.

Key Responsibilities :

- Define and execute the organization's cloud infrastructure and platform engineering strategy.

- Architect and manage highly available, scalable, and secure multi-tenant SaaS platforms.

- Drive platform modernization initiatives and cloud-native transformation programs.

- Establish infrastructure standards, governance frameworks, and operational best practices.

- Collaborate with Engineering, Product, Security, and Architecture teams to align platform capabilities with business objectives.

- Design and optimize AWS-based infrastructure leveraging services such as :

1. EKS (Elastic Kubernetes Service)

2. EC2

3. RDS

4. VPC

5. Route53

6. ELB/ALB

7. IAM

8. CloudWatch

9. S3

10. Lambda

- Build resilient and cost-efficient cloud environments capable of supporting rapid business growth.

- Lead infrastructure capacity planning, disaster recovery, and business continuity initiatives.

- Architect and operate Kubernetes clusters at scale across production and non-production environments.

1. Implement container orchestration strategies for high availability and performance.

2. Establish standards for cluster management, service mesh, ingress controllers, and workload optimization.

3. Improve deployment reliability, scalability, and security of containerized applications.

- Define Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs).

1. Lead reliability engineering initiatives to achieve platform uptime and performance targets.

2. Build proactive monitoring, alerting, observability, and incident management frameworks.

3. Conduct post-incident reviews and drive continuous improvement initiatives.

- Develop and optimize CI/CD pipelines for large-scale SaaS environments.

1. Drive Infrastructure as Code (IaC) adoption using tools such as Terraform, CloudFormation, or Pulumi.

2. Automate provisioning, deployments, configuration management, and operational workflows.

3. Eliminate manual operational overhead through self-service platform capabilities.

- Embed security best practices throughout infrastructure and deployment processes.

1. Implement cloud security controls, secrets management, access governance, and vulnerability management.

2. Support compliance initiatives including SOC2, ISO 27001, GDPR, HIPAA, or similar frameworks.

3. Partner with security teams to strengthen platform security posture.

- Act as the highest level technical authority for DevOps and Platform Engineering.

1. Mentor Senior Engineers, Staff Engineers, and Engineering Managers.

2. Drive engineering excellence, architectural reviews, and technical decision-making.

3. Influence platform roadmap and strategic technology investments across the organization.

Required Skills & Qualifications :

- 10-15 years of experience in DevOps, SRE, Platform Engineering, Cloud Infrastructure, or related domains.

- Experience working in B2B SaaS Product Companies is mandatory.

- Proven experience in roles such as :

1. Principal DevOps Engineer

2. Staff DevOps Engineer

3. Lead DevOps Engineer

4. Principal SRE

5. Staff SRE

6. Platform Engineering Lead

- Demonstrated ownership of organization-wide infrastructure strategy and platform roadmap.

- Expert-level experience with AWS Cloud services :

1. EKS

2. EC2

3. RDS

4. VPC

5. IAM

6. Route53

7. CloudWatch

8. S3

9. Networking Services

- Strong understanding of distributed systems architecture.

- Deep expertise in Kubernetes administration and architecture at scale.

- Strong hands-on experience with :

1. Docker

2. Kubernetes

3. Helm

4. Service Mesh technologies

5. Container Security

- DevOps & Automation

1. Extensive experience with :

i. Terraform

ii. CloudFormation

iii. Ansible

iv. GitOps methodologies

2. CI/CD tools :

i. Jenkins

ii. GitHub Actions

iii. GitLab CI/CD

iv. ArgoCD

v. CircleCI

- Monitoring & Observability

1. Expertise in :

i. Prometheus

ii. Grafana

iii. ELK Stack

iv. Datadog

v. New Relic

vi. OpenTelemetry

- Strong incident management and production support experience.

- Networking & Security

1. Advanced understanding of :

i. VPC Design

ii. Load Balancing

iii. DNS

iv. VPN

v. CDN

vi. Network Security

vii. Zero Trust Architecture

- Experience implementing cloud security best practices and compliance controls.

- Experience supporting large-scale multi-tenant SaaS applications serving enterprise customers.

- Experience building Internal Developer Platforms (IDP).

- Knowledge of FinOps and cloud cost optimization.

- Experience with multi-region deployments and global infrastructure.

- AWS Professional Certifications preferred.

- Exposure to AI/ML infrastructure workloads is a plus.

- Strategic Infrastructure Leadership

- Platform Engineering Excellence

- Cloud-Native Architecture

- Reliability Engineering

- Technical Mentorship

- Cross-Functional Stakeholder Management

- Problem Solving & Decision Making

- Scalability & Performance Optimization

- Security & Compliance Mindset

- 10-15 years of total experience.

- Current or recent experience in a B2B SaaS Product Company.

- Hands-on expertise with AWS, Kubernetes, DevOps, SRE, and Platform Engineering.

- Experience operating multi-tenant SaaS platforms at scale.

- Proven track record of leading infrastructure strategy and platform roadmap at an organizational level.

- Experience in Principal, Staff, or Lead-level engineering roles.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...