Posted on: 15/06/2026
About the Role :
We are looking for a Senior AI Engineer to build a production-grade Document Intelligence Platform powered by AI Agents and Amazon Bedrock. The platform will ingest PDFs of varying quality, extract structured information into predefined schemas, generate field-level confidence scores, and support Human-in-the-Loop (HITL) review workflows.
The ideal candidate should have strong expertise in AI Agents, Document AI, OCR/VLM technologies, RAG architectures, and production-scale LLM applications on AWS.
Key Responsibilities :
- Design and develop AI agents that extract structured data from unstructured and semi-structured documents.
- Build document ingestion pipelines capable of handling low-quality scans, image-based PDFs, and complex layouts.
- Develop confidence-scoring and routing mechanisms to automatically flag low-confidence fields for human review.
- Build and optimize RAG pipelines and domain-specific knowledge bases for grounded and accurate extraction.
- Implement Human-in-the-Loop workflows, reviewer feedback loops, and prompt-injection mitigation controls.
- Develop reprocessing and comparison mechanisms to identify material changes between document versions.
- Optimize model selection, inference costs, latency, and overall platform performance on Amazon Bedrock.
- Mentor engineers and establish reusable standards for extraction and orchestration frameworks.
Required Skills & Experience:
- 6+ years of software engineering experience with strong exposure to AI/ML and Generative AI applications.
- Hands-on experience with AWS and Amazon Bedrock, including prompt engineering, model evaluation, and production optimization.
- Strong experience with AI Agent frameworks such as LangGraph, LangChain, or similar orchestration platforms.
- Expertise in OCR technologies (Amazon Textract or equivalent), Vision Language Models (VLMs), and Document AI solutions.
- Experience building production RAG systems, vector retrieval, chunking, and grounding techniques.
- Proven experience extracting structured data from large PDF documents with high accuracy.
- Strong understanding of confidence estimation, Human-in-the-Loop systems, and evaluation frameworks.
- Experience with Python and production-grade LLM pipelines.
Good to Have :
- Experience with confidence calibration and selective prediction techniques.
- Exposure to Firecracker/microVM sandboxing for untrusted documents.
- Experience with evaluation harnesses, golden datasets, and drift monitoring.
- Knowledge of Commercial Real Estate (CRE) data structures and governance practices.
Did you find something suspicious?