Posted on: 21/09/2026
Platform Engineer (AI Systems) :
The team :
This team builds, operates, and owns the production platform that runs our AI agents. The foundation is a cloud-agnostic, secure-by-default modern CNI Kubernetes stack. What this team ships, this team operates. There is no handoff - the team owns the full product lifecycle, including design, security, packaging, release, and operations.
What you'll work on :
- Build the Kubernetes production stack: ingress, SSO, databases, object storage, caching, secrets, and telemetry, shipped as one versioned package that installs with a single command.
- Extend the self-service layer: custom resources and controllers that auto provisions application teams scoped databases, buckets, caches, and identity realms with least-privilege credentials.
- Build and harden the sandboxed execution services that run untrusted, model-generated code for AI and data workloads: process isolation, resource budgets, durable session workspaces.
- Build the agent runtime: the trusted service that holds the LLM provider connection, streams conversation turns, and dispatches every tool execution into the sandbox rather than running it in-process.
- Write the services and tooling around the platform in Python or Go: HTTP APIs, executors, Kubernetes controllers, bootstrap and release automation.
- Prove the security model: build telemetry, alerting, and automated checks that assert the isolation controls against live deployments.
- On-call responsibilities: team members take turns being on-call for production issues. The rotation is 1-6 (1 week on and 6 weeks off).
Requirements:
- 2-5 years of professional or project experience writing systems software or infrastructure as code.
- Strong fundamentals in algorithms.
- Demonstrated depth in Go or Python.
- Demonstrated depth in Kubernetes or Terraform.
- Academic excellence.
We understand that this is a junior engineering role, and demonstrated depth in one skill from each pair (e.g. Kubernetes and Python) is what passes the interview.
90-day success criteria:
By day 30:
- Run the current platform locally: bootstrap the stack, deploy the sandbox and agent runtime on it, and onboard the example app through the self-service claims.
- Ship a first change to production - small is fine; the point is completing one full design-review-release cycle.
By day 60:
- Own a backlog item end to end: design, implementation, release, and the operational follow-up, with review from the team.
- Land a change in at least two of the three products (stack, sandbox, agent runtime).
By day 90:
- Carry a quarterly backlog item without day-to-day supervision.
- Complete an on-call shadow week and join the 1-6 rotation.
- Run the isolation checks against a live deployment and explain what each control prevents.
Did you find something suspicious?
Posted by
Recruiter
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1673287