HamburgerMenu
hirist

Trianz - Principal AI Architect

Trianz
8 - 12 Years
Bangalore

Posted on: 07/10/2026

Job Description

ABOUT THIS ROLE :

You will own the complete architecture for private and sovereign AI deployment at Trianz.

This is not a cloud-API-consumption role.

You will design the system that runs large language models inside customer environments - choosing the right models, designing the serving topology, setting CPU and GPU routing policies, and ensuring the entire stack is secure, sovereign, and provider-agnostic.

This is a pure IC role with architectural authority over how AI runs in production.

WHAT YOU WILL DO :

- Design the complete private LLM serving architecture : model selection, serving framework (vLLM, TensorRT-LLM, Triton), and runtime topology.

- Define intelligent CPU vs GPU routing policies : which model sizes and prompt types route to which compute tier based on latency, cost, and throughput targets.

- Architect multi-cloud model serving : AWS, Azure, GCP - provider-agnostic, no managed AI lock-in.

- Design on-premises serving for enterprise customers : RedHat OpenShift, VMware - containerised model serving on customer hardware.

- Architect per-tenant model isolation, data-residency compliance, and air-gapped sovereign deployment patterns.

- Define the release architecture for model versions : rollout, staged deployment, rollback, and promotion gates.

- Design the LLM governance framework : model behaviour monitoring, inference audit logging, guardrail architecture.

- Design the DevSecOps pipeline architecture and self-service deployment automation standards.

- Evaluate open-source models (Llama, Mistral, Qwen, Phi) against closed models for specific enterprise use cases.

- Produce architecture sign-off documents and review all AI system designs before implementation.

MUST HAVE :

- Deployed LLMs to production in a real enterprise environment - not just API consumption.

- Hands-on with vLLM, TensorRT-LLM, or Triton in a production serving context.

- Designed CPU cluster inference (Intel Xeon / AMD EPYC) for open-source models.

- Kubernetes at production scale (EKS, AKS, GKE, or OpenShift) - not just local k8s.

- Designed multi-cloud, provider-agnostic AI architectures.

- Experience with model quantization (GPTQ, AWQ, GGUF) and routing trade-offs.

- 8+ years in AI/software architecture with at least 3 years in production LLM systems.

GOOD TO HAVE :

- Experience with RedHat OpenShift for on-premises model serving.

- Familiarity with NVIDIA Dynamo, SGLang, or custom inference schedulers.

- GPU FinOps - cost modelling for mixed CPU/GPU inference fleets.

- Sovereign AI or regulated-industry deployment experience.

- Intel OpenVINO or AMD ROCm for CPU-optimised inference.

Company Overview :

Trianz is an applied AI solutions company that accelerates customer business transformation through AI powered "Transformation Services as a Software Model".

With 25+ years of transforming enterprises, we've evolved to a product-led, platform-driven organization serving global enterprises across Financial Services, Insurance, Healthcare, Hi-Tech, Manufacturing, and other industries.

With global presence across 4 continents, our platform portfolio under the unified Concierto brand delivers end-to-end transformations including solutions for Migrate, Manage, Maximize, Modernize, Insights & Agentic AI, and SecOps delivered through strategic partnerships with leading hyperscalers.

We're building the premier innovation-led organization in the digital transformation space through AI-first methodologies and data-driven excellence RevolutionAIzing Transformations.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...