Posted on: 11/09/2026
What you will do :
- Build multi-agent workflows using LangGraph or similar - routing, state, tool use, human-in-the-loop review.
- Build retrieval systems over domain data: chunking, hybrid search, reranking, and knowledge-graph or ontology-based retrieval where it helps.
- Evaluate before you ship - test harnesses, LLM-as-judge, regression suites - and monitor quality once it is live.
- Handle regulated data properly: de-identification, PHI controls, audit trails, guardrails on inputs and outputs.
- Deploy and run what you build - containers, CI/CD, cloud, safe rollback.
- Work directly with clinical and domain experts to turn unclear workflows into something a system can actually do.
Must have :
- Strong Python and production API experience (FastAPI or similar).
- At least one agentic system you built that real users depend on - you can explain the design, what broke, and what you would change.
- Real RAG engineering, not just a vector store call. Retrieval quality work with numbers behind it.
- Working knowledge of an LLM evaluation and observability stack (LangSmith, Langfuse, Arize, RAGAS or equivalent).
- Docker plus one cloud platform (AWS, Azure or GCP), and comfort owning a service end to end.
- Judgment about when an agent is the wrong answer and a simple pipeline is the right one.
Did you find something suspicious?