HamburgerMenu
hirist

Job Description

Key Responsibilities:

- Design & Build Cloud Platforms: Architect and implement scalable, secure, highly available cloud-native platforms across AWS, GCP, or Azure, including Kubernetes, networking, identity, and multi-tenancy.

- Lead Platform Engineering Initiatives: Own complex, multi-month platform projects end-to-end, from architecture and technical design through implementation, production rollout, and continuous improvement.

- Develop Internal Developer Platforms (IDP): Build and enhance self-service platforms, tooling, and workflows that improve developer productivity, standardisation, and overall developer experience.

- Drive Infrastructure as Code & GitOps: Establish and standardise infrastructure and deployment practices using tools such as Terraform/Pulumi and ArgoCD across engineering teams.

- Build AI/ML Infrastructure: Design and operate infrastructure supporting model serving, inference pipelines, GPU workloads, and LLM integrations, ensuring scalability, reliability, and performance.

- Ensure Security & Compliance: Embed security, governance, and compliance best practices across the platform, aligning infrastructure with frameworks such as SOC 2, ISO 27001, and GDPR.

- Optimise Platform Cost & Performance: Drive FinOps initiatives and optimise cloud/GPU resource utilisation, particularly for large-scale AI/ML and inference workloads.

- Technical Leadership & Documentation: Create design documents and ADRs, communicate architectural decisions clearly, manage technical risks, and collaborate with engineering and business stakeholders to establish long-term platform strategy.

Requirements:

- 5 to 8 years in platform, infrastructure, or SRE with at least 3 years focused on cloud infrastructure or developer platforms.

- Expertise in cloud-native architecture: Kubernetes at scale, managed cloud services, networking, identity federation, and multi-tenancy patterns across AWS, GCP, or Azure.

- Proven ownership of complex, multi-month platform initiatives youve been part of them from whiteboard to production, managing ambiguity and technical risk throughout.

- Proficient in IaC and GitOps fluency Terraform or Pulumi, ArgoCD with experience standardising platform tooling and deployment patterns across engineering teams.

- Hands-on experience with AI/ML infrastructure: model serving, inference pipelines, GPU resource management, or LLM integration patterns.

- Experience in enterprise environments with familiarity with compliance frameworks such as SOC 2, ISO 27001, or GDPR.

- Strong written communication you write clear design docs, maintain useful ADRs, and can explain architectural decisions to both engineers and non-technical stakeholders.

Nice to Have:

- You treat platform as a product and obsess over internal developer experience.

- FinOps or GPU cost optimisation across large inference workloads.

- You default to writing things down and creating shared technical context.

- Building an internal developer platform (IDP) from scratch.

- LLM serving at scale vLLM, Triton, Ray Serve or AI gateway design patterns.

- You push back on short-term thinking and advocate for the right long-term call.

- Open-source contributions to platform or ML infrastructure tooling.

- Youre energised by ambiguity and build clarity where there isnt any.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...