HamburgerMenu
hirist

Senior Technical Staff - AI Platforms/Infra/Applied Research

Scaling Theory Technologies
5 - 10 Years
Bangalore

Posted on: 29/09/2026

Job Description

Job Title : Senior Member of Technical Staff (Senior MTS) - AI Platforms / Infra / Applied Research

Job Description :

We are looking for a Senior Member of Technical Staff (Senior MTS) with approx. 5 years of experience to design, build, and operate high-scale production systems.

In this role, you will own complete systems and workstreams end-to-end - driving architecture decisions, technical trade-offs, and operational excellence from ambiguous requirements to production deployments with minimal scaffolding.

Depending on your expertise demonstrated during the selection process, you will be mapped to one of three specialized tracks :

- Track 1 : AI Platforms - Enterprise AI systems, RAG pipelines, multi-agent frameworks, LLM observability, and guardrails.

- Track 2 : AI Infra - Scalable GPU control planes, fleet management, CUDA, resource isolation, and Kubernetes operators.

- Track 3 : Applied Research - Production-focused ML engineering, fine-tuning/post-training, LLM evaluations, RL/RLHF, and synthetic data pipelines.

Key Responsibilities :

- End-to-End System Ownership : Architect, build, deploy, and operate production-grade services, distributed systems, or infrastructure components with long-term maintainability in mind.

- Track-Specific Technical Execution :

1. AI Platforms : Build enterprise RAG platforms, multi-agent workflows, tool-calling frameworks, vector database integrations, and robust LLM tracing/cost/latency observability.

2. AI Infra : Develop high-throughput GPU control planes, scheduler integrations, K8s controllers/operators, and GPU isolation/sharing systems using CUDA, NVML, or DCGM.

3. Applied Research : Implement production evaluation/benchmarking suites (LLM-as-a-judge), reasoning/agent-planning enhancements, fine-tuning workflows, and data curation pipelines.

- Production Reliability & Operations : Debug complex production issues, maintain operational health, implement retries, fault tolerance, idempotency, monitoring, and drive root-cause resolutions.

- Cross-Functional Leadership : Partner closely with engineering leads, product teams, and enterprise customers to convert customer requirements into robust architectures.

Key Requirements :

- Experience : ~5 years of professional software engineering experience shipping and operating production systems at scale.

- Core Languages : Exceptional programming ability in Python, Go, or Rust.

- Distributed Systems & APIs : Solid command of software architecture, REST/gRPC APIs, databases (PostgreSQL, Redis), and messaging queues (Kafka, NATS, RabbitMQ).

- Cloud & Containerization : Practical hands-on experience with Docker, Kubernetes, and at least one major cloud provider (AWS, Azure, or GCP).

- Engineering Rigor : Strong knowledge of CI/CD, Git, automated testing, logging, tracing (OpenTelemetry, Prometheus, Grafana), and reliability patterns.

- Domain Mastery : Hands-on proficiency in at least one of the core areas : LLM Application Engineering/Frameworks (LangChain, LlamaIndex, CrewAI), GPU/Systems Engineering (CUDA, K8s operators, GPU sharing), or ML/Evaluation Engineering (PyTorch, fine-tuning, RLHF, e-vals).

Desired Attributes :

- Background operating systems in high-growth, fast-paced startup environments.

- Proven ability to make clear, well-articulated technical trade-offs in ambiguous scenarios.

- Strong focus on production impact rather than isolated prototypes or wrappers.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...