Posted on: 09/09/2026
About the Role :
We are a San Francisco-based AI infrastructure company working with leading frontier AI labs to build post-training data and evaluation infrastructure for foundation models. We are hiring a Python Developer to create high-quality datasets, reinforcement learning environments, and benchmarking pipelines used to improve and evaluate state-of-the-art LLMs. This is a remote role with flexible working hours.
Responsibilities :
- Create and curate datasets for LLM post-training (SFT, RLHF, RL, preference optimization).
- Build and maintain RL environments for agent evaluation.
- Develop Python tooling for dataset generation, validation, and transformation.
- Evaluate models on custom benchmarks and testing pipelines.
- Collaborate with research and engineering teams to deliver client-specific post-training datasets.
- Work with terminal-first development workflows and cloud infrastructure.
Required Skills :
- Strong Python programming skills.
- Understanding of LLM fundamentals and post-training concepts (SFT, RLHF, RL).
- Experience working with structured data (JSON, CSV, YAML).
- Git, Linux/Unix command line, and solid software engineering fundamentals.
Good to Have :
- Experience with RAG, agentic AI systems, Hugging Face Transformers, LoRA/PEFT, LangChain or LlamaIndex, vector databases (FAISS, Qdrant, Milvus, Pinecone, Weaviate, ChromaDB), Docker, AWS/GCP, FastAPI/Flask, Bash, CLI tooling, model evaluation frameworks, benchmarking, and AI infrastructure.
Compensation :
- Base Salary : USD $1,250/month
- Equity : ESOP/Equity package included.
- Performance Bonuses : Up to USD $4,000/month (in addition to base salary).
The job is for:
Did you find something suspicious?