Posted on: 06/10/2026
Role Overview :
- Design the complete private LLM serving architecture : model selection, serving framework (vLLM, TensorRT-LLM, Triton), and runtime topology.
- Define intelligent CPU vs GPU routing policies : which model sizes and prompt types route to which compute tier based on latency, cost, and throughput targets.
- Architect multi-cloud model serving : AWS, Azure, GCP - provider-agnostic, no managed AI lock-in.
- Design on-premises serving for enterprise customers : RedHat OpenShift, VMware - containerised model serving on customer hardware.
- Architect per-tenant model isolation, data-residency compliance, and air-gapped sovereign deployment patterns.
- Define the release architecture for model versions : rollout, staged deployment, rollback, and promotion gates.
- Design the LLM governance framework : model behaviour monitoring, inference audit logging, guardrail architecture.
- Design the DevSecOps pipeline architecture and self-service deployment automation standards.
- Evaluate open-source models (Llama, Mistral, Qwen, Phi) against closed models for specific enterprise use cases.
- Produce architecture sign-off documents and review all AI system designs before implementation.
Must Have :
- Deployed LLMs to production in a real enterprise environment - not just API consumption.
- Hands-on with vLLM, TensorRT-LLM, or Triton in a production serving context.
- Designed CPU cluster inference (Intel Xeon / AMD EPYC) for open-source models.
- Kubernetes at production scale (EKS, AKS, GKE, or OpenShift) - not just local k8s.
- Designed multi-cloud, provider-agnostic AI architectures.
- Experience with model quantization (GPTQ, AWQ, GGUF) and routing trade-offs.
- 8+ years in AI/software architecture with at least 3 years in production LLM systems.
Did you find something suspicious?
Posted by
Suhas Balang
Talent Acquisition at Sourcebae
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1676869