Posted on: 07/10/2026
ABOUT THIS ROLE :
You will own the complete architecture for private and sovereign AI deployment at Trianz.
This is not a cloud-API-consumption role.
You will design the system that runs large language models inside customer environments - choosing the right models, designing the serving topology, setting CPU and GPU routing policies, and ensuring the entire stack is secure, sovereign, and provider-agnostic.
This is a pure IC role with architectural authority over how AI runs in production.
WHAT YOU WILL DO :
- Design the complete private LLM serving architecture : model selection, serving framework (vLLM, TensorRT-LLM, Triton), and runtime topology.
- Define intelligent CPU vs GPU routing policies : which model sizes and prompt types route to which compute tier based on latency, cost, and throughput targets.
- Architect multi-cloud model serving : AWS, Azure, GCP - provider-agnostic, no managed AI lock-in.
- Design on-premises serving for enterprise customers : RedHat OpenShift, VMware - containerised model serving on customer hardware.
- Architect per-tenant model isolation, data-residency compliance, and air-gapped sovereign deployment patterns.
- Define the release architecture for model versions : rollout, staged deployment, rollback, and promotion gates.
- Design the LLM governance framework : model behaviour monitoring, inference audit logging, guardrail architecture.
- Design the DevSecOps pipeline architecture and self-service deployment automation standards.
- Evaluate open-source models (Llama, Mistral, Qwen, Phi) against closed models for specific enterprise use cases.
- Produce architecture sign-off documents and review all AI system designs before implementation.
MUST HAVE :
- Deployed LLMs to production in a real enterprise environment - not just API consumption.
- Hands-on with vLLM, TensorRT-LLM, or Triton in a production serving context.
- Designed CPU cluster inference (Intel Xeon / AMD EPYC) for open-source models.
- Kubernetes at production scale (EKS, AKS, GKE, or OpenShift) - not just local k8s.
- Designed multi-cloud, provider-agnostic AI architectures.
- Experience with model quantization (GPTQ, AWQ, GGUF) and routing trade-offs.
- 8+ years in AI/software architecture with at least 3 years in production LLM systems.
GOOD TO HAVE :
- Experience with RedHat OpenShift for on-premises model serving.
- Familiarity with NVIDIA Dynamo, SGLang, or custom inference schedulers.
- GPU FinOps - cost modelling for mixed CPU/GPU inference fleets.
- Sovereign AI or regulated-industry deployment experience.
- Intel OpenVINO or AMD ROCm for CPU-optimised inference.
Company Overview :
Trianz is an applied AI solutions company that accelerates customer business transformation through AI powered "Transformation Services as a Software Model".
With 25+ years of transforming enterprises, we've evolved to a product-led, platform-driven organization serving global enterprises across Financial Services, Insurance, Healthcare, Hi-Tech, Manufacturing, and other industries.
With global presence across 4 continents, our platform portfolio under the unified Concierto brand delivers end-to-end transformations including solutions for Migrate, Manage, Maximize, Modernize, Insights & Agentic AI, and SecOps delivered through strategic partnerships with leading hyperscalers.
We're building the premier innovation-led organization in the digital transformation space through AI-first methodologies and data-driven excellence RevolutionAIzing Transformations.
Did you find something suspicious?