Posted on: 08/06/2026
Job Description :
Role & Responsibilities :
- 10 - 16 years of progressive software engineering experience, with at least 3 - 4 years in an engineering management role leading teams of 10 or more engineers
- Strong programming expertise in one or more of Go, Rust, or Python, with the technical depth to conduct meaningful architecture and code reviews, and to make sound build-versus-buy decisions
- Deep experience designing, building, and operating large-scale distributed systems particularly compute orchestration, job scheduling, container platforms, or infrastructure comparable in complexity to Kubernetes, Borg, Mesos, or similar systems
- Proven track record in platform engineering or software infrastructure - building the foundational systems that other engineers and researchers depend on for development, build, deployment, and production operations
- Strong background in compute-heavy or high-performance computing environments, with working knowledge across the infrastructure stack (compute, storage, networking, GPUs)
- Experience defining and operating against service-level objectives (SLOs) and service-level agreements (SLAs) for mission-critical infrastructure
- Track record of hiring, mentoring, and developing engineers - building cohesive teams that consistently deliver reliable infrastructure in production
- Expertise with relational databases (PostgreSQL, MySQL) and distributed data systems (Kafka, ZooKeeper) in production environments
- Observability and reliability mindset : hands-on experience establishing monitoring, alerting, tracing, on-call rotations, and incident-response practices (Prometheus, OpenTelemetry, or similar)
- Security-conscious approach to infrastructure : familiarity with Linux system hardening, access controls, network segmentation, and compliance-aware development practices
- Strong problem-solving and critical-thinking abilities, with a bias for action and a track record of making pragmatic decisions under ambiguity
- Set the technical direction and architectural standards for the teams portfolio of infrastructure systems, spanning :
i. Compute orchestration and scheduling (job schedulers, service managers, container platforms)
ii. Environment and package management (Nix/Flox ecosystem, reproducible builds)
iii. GPU compute and high-performance computing (CUDA, MPI, performance tuning)
iv. Distributed storage (parallel file systems, SAN, network-attached storage)
v. Networking infrastructure (multicast, core routing, low-latency network engineering)
vi. Linux platform engineering (kernel, eBPF, system performance, provisioning)
vii. Virtualization, cloud infrastructure (GCP), and hybrid compute environments
viii. SDLC tooling, CI/CD, and developer productivity infrastructure
- Lead architecture reviews and design discussions; ensure systems are built for reliability, performance, security, and operability at scale
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Engineering Management
Job Code
1642619