Posted on: 07/07/2026
Job Description :
Role Overview :
As the AI Engineer at 301io, you are not a contributor to an AI team you ARE the AI team. You will set the strategy, make the architectural calls, and be the person everyone looks to when it comes to how AI is built, evaluated, and shipped at this company. You bring the mindset of a full-stack engineer and the depth of an ML practitioner, allowing you to own both the intelligence layer and the systems that surface it to users.
What ownership looks like in this role :
You will decide which models we use, how our agents are designed, how we measure their quality, and how we deploy and iterate on them. This is a greenfield opportunity to build 301ios AI foundation from scratch the tools, the culture, and the standards.
Four Pillars of Ownership :
AI Strategy & Ownership :
Define the AI roadmap, select models and frameworks, and own AI architecture decisions end-to-end at 301io.
Agent Design & Deployment :
Architect, build, and ship production-grade AI agents and multi-agent systems that power real user-facing features.
Evaluation & Quality :
Design and operate eval pipelines that measure model accuracy, safety, and regression making AI reliable, not just impressive.
Full-Stack AI Integration :
Bridge model capabilities with product surfaces, owning the API layer, frontend integrations, and infrastructure that brings AI to life.
Key Responsibilities :
AI Strategy & Architecture :
- Own the companys AI roadmap model selection, framework decisions, build vs. buy trade-offs
- Evaluate and select foundation models (OpenAI, Anthropic, Google, open-source) for each use case
- Design scalable AI architectures that support both experimental and production workloads
- Stay ahead of the curve on emerging AI capabilities, papers, and tooling and bring insights to the team
- Define AI engineering standards, best practices, and documentation across the org
LLM Integration & Prompt Engineering :
- Design, version, and maintain prompt systems for production LLM applications
- Implement advanced prompting techniques chain-of-thought, few-shot, RAG, structured outputs, tool use
- Build robust context management systems memory, compression, retrieval, conversation state
- Optimize LLM calls for latency, cost, and quality across all product surfaces
- Manage model versioning and graceful migration between model versions
AI Agent Design & Deployment :
- Architect and build autonomous AI agents and multi-agent orchestration systems
- Implement agentic patterns ReAct, Plan-and-Execute, Reflexion, tool-using agents, long-horizon tasks
- Define agent boundaries, safety constraints, and human-in-the-loop intervention points
- Build production-ready agent pipelines using frameworks such as LangGraph, CrewAI, AutoGen, or custom orchestration
- Monitor agent behavior in production and iterate rapidly on failures
Retrieval-Augmented Generation (RAG) & Knowledge Systems :
- Design and implement RAG pipelines chunking strategies, embedding models, vector stores, reranking
- Select and manage vector databases Pinecone, Weaviate, Qdrant, pgvector
- Build hybrid retrieval systems combining dense and sparse search for optimal recall and precision
- Implement knowledge graph integrations and structured data retrieval where appropriate
- Continuously improve retrieval quality using evaluation metrics and A/B testing
Evals, Quality, & Reliability :
- Design and operate comprehensive evaluation frameworks for LLM outputs and agent behavior
- Build automated eval pipelines unit evals, regression suites, LLM-as-judge, human preference evals
- Define quality metrics and SLOs for AI features accuracy, hallucination rate, task completion, latency
- Implement CI/CD for AI eval-gated deployment that prevents quality regressions from reaching production
- Manage eval datasets, annotation workflows, and ground-truth curation
- Integrate evaluation tools such as Braintrust, LangSmith, PromptFoo, or Weights & Biases
ML Model Training, Fine-Tuning & Deployment :
- Fine-tune and adapt foundation models using techniques like LoRA, QLoRA, PEFT, RLHF, and DPO
- Manage ML training workflows and experiment tracking MLflow, W&B, Comet
- Package and deploy models using serving infrastructure vLLM, TGI, Triton, Ray Serve, BentoML
- Optimize model inference for production quantization, batching, caching, speculative decoding
- Implement model monitoring for drift, degradation, and safety incidents
AI Infrastructure & MLOps :
- Build and maintain ML pipelines for data processing, training, and deployment
- Manage GPU/TPU compute resources efficiently across cloud and on-prem environments
- Implement feature stores, model registries, and experiment tracking infrastructure
- Containerize and orchestrate ML workloads using Docker, Kubernetes, and cloud ML platforms
- Collaborate with DevOps on CI/CD integration, autoscaling, and cost optimization for AI workloads
Full-Stack AI Integration :
- Build APIs and backend services that expose AI capabilities to product surfaces
- Develop frontend components that deliver intelligent, AI-powered user experiences
- Implement streaming, real-time inference, and async processing patterns for responsive AI UX
- Integrate AI features into existing product workflows with minimal friction
- Own the end-to-end stack: from model weights to the UI element the user sees
AI Safety, Ethics & Governance :
- Implement guardrails, content filtering, and output validation for all production AI systems
- Define and enforce responsible AI principles across model selection and deployment
- Conduct red-teaming and adversarial testing to identify failure modes before launch
- Maintain data privacy compliance PII handling, data retention, regulatory requirements
- Document model cards, system cards, and decision logs for governance and auditability
Requirements :
Experience :
- 6+ years of total software engineering experience
- 3+ years of hands-on experience deploying AI/ML models and systems in production
- Demonstrated experience building and shipping AI agents or agentic systems
- Track record of owning AI features end-to-end, from research to production
- Prior full-stack engineering experience proficiency in both backend and frontend domains
Core AI/ML Skills :
- LLM APIs: OpenAI, Anthropic, Google Gemini
- Agentic frameworks: LangChain, LangGraph, CrewAI
- RAG pipeline design and vector databases
- Prompt engineering and structured outputs
- Evaluation frameworks and LLM-as-judge
- Fine-tuning: LoRA, QLoRA, PEFT, RLHF, DPO
- ML experiment tracking: W&B, MLflow, Comet
- Model serving: vLLM, TGI, Triton, BentoML
- Embedding models and retrieval systems
- AI safety, guardrails, and red-teaming
Full-Stack Engineering Skills :
- Python (primary language for AI work)
- REST APIs / GraphQL / streaming (SSE, WebSockets)
- React or equivalent modern frontend framework
- SQL and NoSQL databases
- Docker & Kubernetes for ML workloads
- Cloud: AWS / GCP / Azure ML services
- CI/CD integration for model deployment
- Git, code review, and engineering best practices
Nice-to-Have :
- Experience with multimodal models vision, audio, code generation
- Contributions to open-source AI projects or published research
- Experience with on-prem or edge model deployment
- Knowledge of compiler-level optimizations for inference TensorRT, ONNX, MLIR
- Background in cognitive architectures or AI planning systems
- Experience with knowledge graphs and structured reasoning
Mindset & Soft Skills :
- First-principles thinker who questions assumptions and reaches for root causes
- Comfortable operating in ambiguity and making principled decisions with incomplete information
- Ownership mentality you ship things, you monitor them, and you take responsibility for outcomes
- Excellent communicator who can explain AI trade-offs to engineers, PMs, and executives
- Relentlessly curious about new AI developments and fast to experiment with new capabilities
- Passionate about quality not just making AI work, but making it reliable and trustworthy
Did you find something suspicious?