HamburgerMenu
hirist

Staff Product Development Engineer - AI Powered System

Careernet
8 - 16 Years
Bangalore

Posted on: 25/06/2026

Job Description

Job Summary :

We are hiring a Staff Product Development Engineer to serve as the technical lead for Model Evaluation & Systems within Foundations Engineering AI Production Engineering.

This role defines and owns the measurable quality standards for AI-powered systems across Ad Platforms. You will architect structured evaluation pipelines, automated regression detection systems, benchmarking frameworks, and model promotion criteria (Eval ? Beta ? GA).

You will partner closely with AI Core Engineering, Product, Security, Architecture, and Governance teams to ensure AI models meet clearly defined quality, safety, and performance thresholds before scaling externally or impacting revenue workflows.

This role blends system design, AI evaluation rigor, and cross-functional technical influence. You will define what production-ready means for AI models across the organization.

Responsibilities :

Evaluation Architecture & Standards :

- Define structured evaluation frameworks for AI-powered systems across Ad Platforms.

- Architect automated regression detection pipelines for model updates and provider changes.

- Establish benchmarking standards across multiple model providers and versions.

- Define measurable quality thresholds for model promotion (Eval, Beta, GA).

- Drive implementation of reproducible, engineering-owned evaluation workflows integrated into release processes.

Model Quality & Engineering Governance :

- Partner with AI Core Engineering leadership and Governance stakeholders to define safety, response consistency, and hallucination scoring standards.

- Design prompt and workflow drift detection mechanisms for long-running AI systems.

- Ensure evaluation processes generate defensible, engineering-driven metrics for executive readiness reviews.

- Establish documentation and certification standards for model readiness and promotion.

- Formalize engineering-controlled quality gates for AI releases.

Engineering & Automation :

- Build scalable evaluation pipelines integrated with CI/CD workflows and release gating systems.

- Design automated A/B testing systems for model comparison and provider validation.

- Define and maintain structured evaluation datasets and test harnesses.

- Partner with Runtime & Reliability engineers to align evaluation benchmarks with real-world production behavior.

- Continuously refine evaluation infrastructure to improve repeatability, traceability, and coverage.

Cross-Functional Engineering Influence :

- Serve as the technical authority for AI quality and evaluation standards within Foundations Engineering.

- Influence model selection and release decisions in partnership with AI Core and Architecture leadership.

- Mentor engineers on evaluation methodologies and reproducible testing systems.

- Provide engineering-driven recommendations during model readiness and production certification reviews.

Basic Qualifications :

- 8+ years of backend or systems engineering experience.

- Strong proficiency in Python.

- Experience designing automated testing or evaluation frameworks at scale.

- Experience operating AI-powered or model-dependent systems in production.

- Strong understanding of CI/CD integration and release gating processes.

- Experience defining measurable quality standards across distributed systems.

- Proven ability to influence architectural and engineering standards across multiple teams.

Preferred Qualifications :

- Experience designing benchmarking systems for LLMs or ML models.

- Familiarity with hallucination detection, safety scoring, or response consistency analysis.

- Experience building A/B testing infrastructure for model comparison.

- Exposure to multi-provider model environments (Azure, OpenAI, Bedrock).

- Background in experimentation systems or engineering-led release governance.

Experience with :

- Designing automated regression testing frameworks for AI systems.

- Building structured evaluation datasets and benchmarking pipelines.

- Model comparison systems and A/B testing infrastructure.

- Prompt drift detection and workflow validation.

- CI/CD integration for automated evaluation gating.

- AI provider APIs and model performance characteristics.

- Observability and analytics systems for quality measurement.

- Cross-functional collaboration with AI Core, Architecture, and Governance leadership.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...