Posted on: 07/08/2026
Role Summary :
You will own Mynts infrastructure and delivery lifecycle end-to-end: architect and provision our AWS footprint (EC2, ECS, Lambda, RDS, S3, ECR, Route 53), containerise and orchestrate services with Docker and Kubernetes, and run the data-and-messaging backbone (PostgreSQL, Redis, Kafka).
Youll build the CI/CD pipelines that ship our Node/React/Python stack safely, tame the runtime and web tier (NVM, Pyenv, nginx/Apache, Caddy/Traefik, Certbot), and keep every environment Dev, QA, Staging, and Production fast, observable, and cheap to run.
Above all, youll bring seasoned, opinionated judgement from multiple cloud ecosystems to recommend the most cost-effective, seamlessly-scalable path for every decision.
The AI-Augmented Edge :
At Pyvot, AI is your primary workforce. Use Claude Code, Cursor, and Gemini to draft IaC, generate pipeline configs, review infra diffs, and reason about failure modes targeting a 3x5x efficiency gain.
Practise Cross-LLM Validation: one model proposes the architecture, another stress-tests it for cost, blast-radius, and scaling limits. Your value is measured by uptime, unit economics, and how gracefully the platform scales not tickets closed.
Core Responsibilities :
Cloud Architecture & Infrastructure-as-Code : Be the Architect :
- Multi-Service AWS Footprint : Design, provision, and harden EC2, ECS, Lambda, RDS (PostgreSQL), S3, ECR, and Route 53 as reproducible infrastructure not hand-crafted snowflakes.
Reliability, Observability & Site Reliability Engineering :
- Monitoring & Alerting : Stand up metrics, logs, and traces (CloudWatch, Prometheus / Grafana, ELK, PagerDuty) with actionable alarms not alert fatigue.
- Availability & Recovery : Define and defend SLOs, RTO / RPO targets, backups, and disaster-recovery / failover drills you have actually rehearsed.
- Incident Response : Be first responder for infrastructure incidents diagnose, mitigate, and run blameless post-mortems that harden the system.
- Boot & Persistence Hygiene : Ensure services survive reboots, scaling events, and deploys (pm2 / systemd / docker restart policies) with no silent drift.
Advisory, Standards & Team Leadership : Be the Trusted Voice :
- Architecture Recommendations : Proactively bring cost / scaling / reliability trade-off proposals to the CTO youre hired for judgement, not just execution.
- Runbooks & Standards : Maintain infrastructure runbooks, deployment standards, and on-call playbooks the whole team can follow.
- Mentoring : Level up engineers on deployment, containers, and operational best-practice; make the platform something anyone can safely ship to.
- Security-Aware Ops : Partner with the CyberSecurity Lead on hardening, IAM, network segmentation, and audit-friendly, deny-delete logging.
Roles Requirements : Intersection of Cloud, Delivery & Reliability
Were looking for a full-spectrum infrastructure engineer equal parts cloud architect, platform / DevOps engineer, and reliability practitioner: someone at ease with a Terraform diff, a Kubernetes manifest, a Postgres failover, and a 2 AM scaling alert and seasoned enough to tell us the smartest, cheapest way to run all of it.
Must-Have Skills & Experience :
- DevOps / Cloud Engineering Experience : 6 - 10 years running production infrastructure for a SaaS / cloud-native product end-to-end.
- AWS Depth : Hands-on with EC2, ECS, Lambda, RDS, S3, ECR, Route 53, IAM, and VPC networking in real production.
- Multi-Cloud Exposure : Genuine experience across more than one cloud / ecosystem (AWS plus GCP / Azure / DO / etc.) able to compare and choose, not just operate one.
- Containers & Orchestration : Strong Docker and Kubernetes (EKS or self-managed) autoscaling, rollouts, and resource management.
- IaC & CI/CD : Terraform / CloudFormation and pipeline ownership (GitHub Actions or equivalent) with automated, rollback-safe deploys.
- Data & Runtime Ops : PostgreSQL operations (backup / replication / tuning), Redis, and the web / runtime tier (nginx / Apache, Caddy / Traefik, Certbot, NVM, Pyenv).
- Cost & Scale Judgement : A track record of cutting cloud spend and scaling systems smoothly with the numbers to prove it.
Should-Have Skills & Experience :
- Event Streaming : Production experience with Kafka (or equivalent) for high-throughput, ordered event pipelines.
- Observability & SRE : CloudWatch / Prometheus / Grafana / ELK / PagerDuty, SLOs, RTO / RPO, and rehearsed DR / failover.
- Networking & TLS : Solid grasp of DNS, load balancing, reverse proxies, and certificate automation.
- Scripting & Automation : Comfortable in Bash plus Python / Node to automate anything that repeats.
- Agile & Cross-Functional Collaboration : Comfortable in Agile / Scrum delivery and working closely with engineering, product, and leadership.
Good-To-Have Skills & Experience :
- Education & Certifications : Degree from a premier institute (IITs / NITs / BITS / IIITs) and/or AWS Solutions Architect / DevOps / CKA / CKAD a strong plus.
- Security-Adjacent Ops : Familiarity with IAM hardening, secrets management, and compliance friendly logging (works alongside our Security Lead).
- AI-Augmented Tooling : Experience using AI / LLM tools (Claude Code, Cursor, Gemini) for IaC drafting, config review, and incident reasoning.
- Startup / Scale-Up Experience : Prior ownership of infrastructure through rapid growth on a lean budget.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1661238