Abridge06/22/2026

Product Lead, AI/ML (Evals)

Salary not specified
MARKET
15,900median for this role
Data Scientist · 88 jobs with disclosed pay
5,800half of the offers: 12,695–20,77743,793
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьУдалёнка

Responsibilities

  • 01Drive product strategy and execution for the evals platform. Own the roadmap across the eval lifecycle and own outcomes against it.
  • 02Build the shared measurement infrastructure. Help build the systems that let any pod define quality, run experiments, compare models, and watch production. Own the standards for LLM judges, rule-based evaluators, human annotation, and online monitoring, and be clear about where each belongs.
  • 03Make model selection fast and routine. Give teams a repeatable way to evaluate a new frontier model within days of release, and a defensible framework for when a post-trained specialty model is worth it over a prompted frontier one. Keep the eval system model-agnostic so it stays a neutral referee.
  • 04Own the operating model across pods. Define the eval gates from early build through GA and steady-state monitoring, which gates are hard versus advisory, and who owns non-negotiable floors like critical-error rates. Land this across pods you don't own, without formal authority.
  • 05Cross functional execution. Work with Data Engineering on the de-identification pipeline, Data Science on bootstrapping judge quality with less human annotation, Clinical Science on flagged production cases, and the agent platform team as workflows go agentic.
  • 06Operate with a high bar for quality, speed, and accountability.

Requirements

  • 015+ years of product management experience with significant ownership of ML powered products or platform systems.
  • 02Deep understanding of how to measure and improve model quality, including evaluation frameworks, annotation pipelines, and benchmark design.
  • 03Strong technical fluency across ML, data pipelines, and distributed systems.
  • 04Experience working closely with ML researchers and engineers to drive impact in production.
  • 05Ability to balance long term architectural investments with near term quality improvements.
  • 06Strong communication skills and the ability to translate complex technical concepts into clear decisions and narratives.
  • 07A track record of delivering high quality products in domains where accuracy, reliability, and trust are paramount.

What we offer

  • 01Flexible work hours
  • 02Inclusive culture
  • 03Ongoing learning opportunities
  • 04Opportunity to work in a fast-paced, high-growth startup
  • 05Offices located in San Francisco, New York, and Pittsburgh