HamburgerMenu
hirist

Job Description

About the Role :

Embed with an assigned Product Engineering group as a functional team member; participate daily in standups, design reviews, backlog grooming, sprint planning, and release readiness calls.

Key Responsibilities :

- Build and maintain deep product knowledge - what the service does, who uses it, how pieces fit together - so platform guidance is grounded in product context, not generic best practice.

- Review pull requests and architecture proposals in the flow of work, using the lens : 'Will this run reliably, scale safely, and remain secure under 24/7 SRE operations?'

- Trace data flows and troubleshoot production issues end-to-end across application and infrastructure boundaries; diagnose and fix problems without escalating to core product developers.

- Guide teams through migration onto golden paths and the internal developer platform; identify and remove friction that prevents self-service deployments.

- Bring services to operational maturity before handoff to SRE - establish monitoring, alerting, runbooks, failover strategies, and partner with SRE on validation.

- Conduct hands-on code review and patch application code for security remediation, version upgrades, and performance optimization.

- Own the monthly patching cycle for the assigned group : OS, runtime, container base image, and dependency patches applied, functionally tested, and regression-verified.

- Run scheduled performance, load, and soak testing against reference workloads; triage findings and implement remediation - pull core engineering only where product changes are required.

- Join major incidents alongside SRE; lead root-cause analysis; feed platform and architecture gaps back into the roadmap.

- Champion developer success : workshops, pair programming, and an escalation path when the platform doesn't fit a use case.

- Partner with Release Engineering on deployment strategies; with SRE on operational standards; and with InfoSec on security controls built into the workflow.

Must Have :

- 5 - 8 years in software engineering, DevOps, SRE, or platform engineering with substantial time writing and debugging production code.

- Track record of working directly inside product engineering teams (not through ticket queues); you've been embedded and know how to be a team member.

- Ability to diagnose and fix production issues end-to-end across application & infrastructure.

- Build-capable in Go (Golang) and Python; able to read, debug, and patch Node.js/TypeScript, Dart/Flutter, Java, .NET, C++, and C code.

- SQL expertise (TSQL, PLpgSQL) for performance tuning and migration pipelines.

- Hands-on experience with Amazon EKS, Kubernetes, Cilium, Helm & related cloud-native technologies for deploying, securing, and operating production workloads at scale. Working knowledge of GCP/GKE.

- Experience with API gateways and service-mesh technologies, including Kong, Envoy, and Istio, with a focus on secure API management, traffic routing, and service-to-service communication.

- Proficiency with Terraform, Terragrunt, and Infrastructure-as-Code; ability to review and improve IaC.

- CI/CD pipeline design and implementation (GitHub Actions for build, Harness for deploy); understanding of artifact caching, quality gates, security scanning.

- Expertise in integration and end-to-end testing (Playwright, Cypress, PyTest, Testcontainers); ability to guide teams on test strategy.

- Performance testing (K6, Locust, JMeter); capacity planning and scalability analysis.

- Database knowledge (DynamoDB, PostgreSQL, MySQL, MongoDB, Cassandra, SQL Server); query optimization and data modeling review.

- Message streaming and event-driven systems (Confluent Kafka, Redis); distributed systems concepts.

- Observability fundamentals : OpenTelemetry instrumentation, metrics (Prometheus, Grafana), logs, distributed tracing, and alerting strategies.

- Experience with Databricks, Databricks Apps, and dashboards, including building and supporting data applications and visualizations.

- Security and supply chain practices : SonarQube, Dependabot, JFrog; ability to drive remediation.

- Daily use of AI-assisted development tools (Claude Code, GitHub Copilot, CodeRabbit).

Nice to Have :

- Mobile DevOps experience (Fastlane with Flutter/Dart, Swift, Kotlin; automated mobile UI testing).

- Backstage-based internal developer platform experience.

- Familiarity with internal frameworks (Gofr, NodeFr/Zode).

- LLM integration and prompt engineering skills.

- Cloud architecture certifications (AWS Solutions Architect, GCP Associate Engineer).

- Kubernetes certifications (CKA, CKAD).

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...