Posted on: 01/05/2026
We are seeking a technically deep and people-first AI Engineering Manager to lead the design, delivery, and responsible governance of enterprise-grade AI systems. You will own cross-functional teams building LLM applications, agentic workflows, multimodal pipelines, and ML platforms - while driving a culture of eval-first engineering, cost discipline, and production reliability at scale.
Key Responsibilities :
1. Technical Architecture & GenAI :
- Architect LLM applications across OpenAI, Claude, Gemini, and open-source models
- Design and govern MCP server/client architectures and agentic tool registries
- Build hybrid inference pipelines routing tasks to reasoning vs. fast models
- Manage thinking token budgets, verifier-generator patterns, and chain-of-thought evaluation
- Architect multimodal pipelines - vision-language, speech, document AI
- Optimize RAG pipelines - vector DBs, embedding strategies, retrieval tuning
- Lead prompt engineering, RAG evaluation (BERTScore, LLM-as-judge), and hallucination tracking
- Implement guardrails, safety filters, and compliance frameworks
2. Agentic Systems & Memory :
- Lead multi-agent orchestration using LangGraph, AutoGen, and MCP-native patterns
- Design long-term agent memory - episodic, semantic, and procedural layers
- Define human-in-the-loop controls and safety boundaries for autonomous agents
- Establish agent interoperability standards across internal and third-party systems
- Own security posture for agentic systems - prompt injection, tool misuse, scope creep
3. MLOps & Platform Engineering :
- Establish CI/CD for ML models - Argo, GitHub Actions, Tekton
- Scalable inference with Kubernetes, vLLM, Ray, Triton
- Model optimization - quantization, LoRA fine-tuning, distillation for production
- Evaluate and deploy SLMs for edge / on-device use cases
- Experiment tracking - MLflow, Weights & Biases; feature stores with Feast
- Observability - OpenTelemetry, Prometheus, Grafana; drift and bias monitoring
- Oversee synthetic data generation pipelines and data flywheel strategy
4. Eval-First Engineering Culture :
- Champion eval-driven development as a hard deployment gate
- Build team capability to write domain-specific evaluation suites
- Track regression benchmarks across model versions and prompt changes
- RAG evaluation using BERTScore, LLM-as-judge, RAGAS frameworks
- Implement AIOps - prompt caching, semantic caching, token budget governance
- Define cost-per-inference targets and track AI infrastructure spend
5. Team Leadership & Delivery :
- Lead and mentor AI/ML, Data Engineering, and MLOps engineers
- Own end-to-end delivery from requirements through production monitoring
- Drive Agile, DevOps, and MLOps practices across the team
- Govern AI coding assistant adoption and measure developer productivity lift
- Conduct code and design reviews; enforce architectural standards
- Build a high-performance, psychologically safe engineering culture
6. Governance & Responsible AI :
- Maintain EU AI Act documentation - Articles 11 & 13 conformity assessments
- Conduct high-risk AI assessments and manage incident reporting obligations
- Align with NIST AI RMF 2.0 and ISO/IEC 42001
- Implement responsible AI - fairness, explainability, privacy controls
- Translate business requirements into AI solutions for non-technical stakeholders
- Partner with Product, Architecture, Legal, and Security
Experience Required :
- 10+ Years in software / data / AI 5+ Years in AI/ML or GenAI systems 3+ Years leading engineering teams 1+ Years with agentic / LLM-native production
Required qualifications :
- Bachelor's / Master's in Computer Science, AI, Data Science, or equivalent
- Hands-on with OpenAI, Anthropic, Azure OpenAI, or open-source LLMs in production
- Proficiency in Python; familiarity with Go or TypeScript a plus
- Deep understanding of RAG, vector databases, and embedding strategies
- MCP architecture and agentic tool-calling standards
- Reasoning model operations and hybrid model routing (fast vs. thinking models)
- Multimodal pipeline experience - vision, audio, document AI
- MLOps tooling - MLflow, W&B, Argo, Feast
- EU AI Act literacy and NIST AI RMF 2.0 / ISO 42001 familiarity
- Kubernetes and distributed inference (vLLM, Ray, Triton)
- Eval frameworks - Braintrust, LangSmith, Inspect, or equivalent
- AIOps - prompt caching, semantic caching, token budgeting
Preferred / bonus :
- Managed AI coding assistant governance across an engineering org
- Published evals, benchmarks, or open-source AI tooling
- Experience leading AI red-teaming or safety functions
- AI incident response and post-mortems (hallucination events, agent misuse)
- Synthetic data generation and data flywheel design experience
- ISO/IEC 42001 implementation or audit experience
Technical Skills :
- Python / TypeScript MCP architecture & agentic tool registries
- LangGraph / AutoGen Reasoning models (o3, Claude, DeepSeek R1)
- OpenAI / Azure OpenAI / Anthropic Claude Multimodal AI (vision, audio, document)
- RAG pipelines / Vector DBs
- Kubernetes / vLLM / Ray / Triton Agent memory architecture (episodic, semantic)
- LoRA / QLoRA / quantization
- MLflow / Weights & Biases Braintrust / LangSmith / Inspect (evals)
- Argo / GitHub Actions / Tekton AIOps / semantic & prompt caching
- OpenTelemetry / Prometheus / Grafana Synthetic data pipelines / data flywheel
- Feature store (Feast) EU AI Act / NIST AI RMF 2.0 / ISO 42001
Did you find something suspicious?