HamburgerMenu
hirist

Algoleap Technologies - MLOps/LLM Infrastructure Engineer

AlgoLeap Technologies
7 - 10 Years
Multiple Locations

Posted on: 30/09/2026

Job Description

Role Overview:

Responsible for hosting, deploying, and operating open-weight LLMs within a sovereign cloud environment. The role focuses on GPU infrastructure, model serving, performance optimization, and reliable model lifecycle management.

Key Responsibilities:

- Deploy and operate models such as GPT, LLaMA, Gemma, Mistral, and other product/open-weight models within sovereign cloud.

- Manage GPU provisioning, capacity planning, utilization, and performance optimization.

- Implement LLM inference serving using platforms such as vLLM, Triton, or similar frameworks.

- Build model deployment, versioning, rollback, and lifecycle management processes.

- Develop MLOps pipelines for model packaging, testing, deployment, and monitoring.

- Monitor latency, throughput, GPU utilization, availability, and inference costs.

- Implement scalable and highly available model-serving infrastructure using Kubernetes and containers.

- Work closely with platform, security, and gateway teams to ensure secure model access and governance.

- Troubleshoot production issues across GPU, inference, Kubernetes, networking, and model-serving layers.

Tech Stack:

- MLOps, LLM infrastructure, model serving, GPU-based inference, Kubernetes, vLLM, NVIDIA Triton, TensorRT-LLM, LLaMA, Gemma, Mistral, GPT, Docker, CI/CD, model registries, observability, quantization, batching, caching, GPU memory management, inference optimization, private/sovereign-cloud environments.

Preferred Skills:

- NVIDIA GPUs, CUDA, Helm, Prometheus/Grafana, MLflow, automated model deployment pipelines.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...