Posted on: 09/10/2026
About the role :
We are looking for an SDE III - Agentic AI to design, build and own production-grade LLM agents that solve complex healthcare problems. This is not about building wrappers around an LLM or experimenting with agents through POCs. You'll work on systems where LLMs reason across multiple steps, use tools, interact with backend systems and make decisions inside real workflows. You will own these systems end-to-end from agent and backend design to evals, deployment, observability and post-production improvement.
Roles & Responsibilities :
- Design, build and own production-grade LLM agents used in real healthcare workflows.
- Build the backend services, APIs and tools that agents interact with.
- Design agent workflows including tool calling, orchestration, context, memory, guardrails and failure handling.
- Build rigorous evaluation systems using golden datasets, deterministic checks, LLM-as-judge, human review and trajectory evaluation.
- Define and track measurable quality metrics such as precision/recall/F1, task success, tool-call accuracy and regression rates.
- Build eval-gated release processes including automated regression suites, shadow/canary deployments and rollback mechanisms.
- Instrument and monitor agents in production across quality, reliability, latency and cost.
- Debug failures across prompts, models, tools and underlying backend services.
- Make pragmatic decisions around model selection based on accuracy, cost and latency.
- Know when an agent is the right solution - and when deterministic software is better.
Requirements :
- 6+ years of software engineering experience, with strong ownership of production backend systems.
- 1+ year of experience building LLM agents in production with real users.
- Experience owning at least one agent end-to-end - design, prompts, tools, evals, deployment and post-launch improvement.
- Strong backend engineering and distributed systems fundamentals.
- Hands-on experience evaluating AI systems against real datasets and clearly defined metrics.
- Strong understanding of APIs, production architecture, reliability and observability.
- Ability to reason about model choice, prompting, tool design, context and failure modes.
- Strong ownership and ability to work across engineering, product and domain teams.
- Exceptional total compensation with significant equity upside.
Did you find something suspicious?