HamburgerMenu
hirist

Senior AI Infrastructure/AI Platform Engineer

Futureleap Search Partners
5 - 15 Years
Multiple Locations

Posted on: 29/06/2026

Job Description

Senior AI Platform Engineer

The Role :

We are hiring a Senior AI Platform Engineer to build the platform layer that operationalizes core AI systems.

This role is centered on R&D-to-platform rollout: taking technically sophisticated model systems and making them usable, reliable, and scalable inside the product stack. You will work across model pipelines, training systems, inference infrastructure, distributed execution, and backend platform architecture.

This is not a thin integration role. It requires strong engineering depth and enough model understanding to work effectively with systems involving fine-tuning, reinforcement learning (RL), alignment, interpretability, agent execution, and inference optimization across LLMs, agents, and tabular foundation models.

Responsibilities :

- Own the platformization of internal AI libraries, turning research-heavy systems into robust platform capabilities with stable APIs, execution layers, observability, and deployment paths.

- Build and scale training and post-training infrastructure for workflows including supervised fine-tuning (SFT), reinforcement learning (RL), evaluation, model adaptation, and agent optimization.

- Design the integration layer between research systems and product infrastructure, including job orchestration, artifact management, dataset versioning, experiment lineage, and runtime control surfaces.

- Build inference systems that support complex model behaviors under production constraints, including latency, throughput, cost efficiency, debuggability, and safety.

- Design for multi-cluster and distributed execution, including scheduling, fault tolerance, checkpointing, retries, workload isolation, and heterogeneous compute environments.

- Operationalize internal fine-tuning, reinforcement learning, interpretability, tracing, and behavioral inspection systems into production-ready platform capabilities.

- Build common platform primitives for model lifecycle management across training, evaluation, serving, rollback, and monitoring.

- Partner closely with research teams to translate model-science complexity into production architecture without sacrificing technical rigor.

- Improve platform reliability for long-running and failure-prone AI workloads, especially where model behavior, system behavior, and infrastructure behavior interact in complex ways.

- Ensure alignment, interpretability, and auditability are embedded into system design, particularly for enterprise and regulated deployments where model outputs and decisions must be explainable.

Example Problems You Might Work On :

- Turn internal fine-tuning and reinforcement learning frameworks into production-grade services supporting multiple model families, datasets, and evaluation loops.

- Build rollout infrastructure for new model-science capabilities so research systems can be safely and incrementally exposed inside the platform.

- Integrate model tracing and interpretability capabilities into training and inference pipelines to enable debugging and behavioral inspection.

- Build inference architecture for large language models and agent systems that balance cost, performance, explainability, and runtime control.

- Design distributed execution flows across clusters for long-running training, evaluation, and analysis workloads with strong guarantees around recovery and reproducibility.

- Unify workflows across LLMs, agents, and tabular models without forcing a one-size-fits-all abstraction.

- Build platform interfaces that allow downstream teams to launch, inspect, evaluate, and deploy complex model workflows without rebuilding research infrastructure.

Requirements :

- Strong experience building and shipping complex AI/ML systems in production.

- Deep backend and platform engineering experience, especially in Python, distributed services, workflow orchestration, data systems, and cloud infrastructure.

- Hands-on experience with one or more of :

1. Fine-tuning systems

2. Reinforcement learning pipelines

3. Inference infrastructure

4. Distributed training

5. Model serving

6. Evaluation systems

- Strong understanding of the systems implications of modern model workflows across LLMs, agents, and structured/tabular foundation models.

- Experience scaling workloads across clusters and production environments with strong instincts around reliability, observability, and performance.

- Ability to work across research code, systems code, and product infrastructure without losing rigor at either layer.

- Strong technical judgment around trade-offs between model quality, infrastructure complexity, scalability, interpretability, and operational cost.

Strong Bonus Signals :

- Experience with alignment, interpretability, or AI safety systems.

- Experience with multi-cluster scheduling, inference optimization, or serving infrastructure for large models.

- Experience converting internal research frameworks into reusable platform capabilities.

- Experience debugging production failures caused by interactions between model behavior, orchestration systems, and infrastructure.

- Experience with agent runtimes, tool orchestration, long-horizon execution, or stateful model systems.

What This Role Is Not :

- Not a prompt engineering role.

- Not a glue-code integration role.

- Not a research-only role disconnected from deployment.

- Not a junior role.

This role is for engineers who want to work at the intersection of model science, platform architecture, and production systems, building the infrastructure that enables advanced AI capabilities to operate reliably at scale.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...