We are looking for an LLM Trainer (Generalist) to evaluate, review, and improve the quality of responses generated by Large Language Models (LLMs). You will assess AI outputs across multiple domains such as reasoning, writing, factual accuracy, coding (basic), mathematics, and safety.
Your evaluations and annotations will directly contribute to training next-generation AI models through high-quality human feedback (RLHF/SFT), helping make AI systems more accurate, reliable, and aligned with human expectations.
Key Responsibilities :
- Evaluate AI-generated responses for correctness, completeness, relevance, and clarity.
- Compare multiple LLM responses and identify the best-performing answer.
- Apply annotation guidelines consistently across diverse tasks.