Mercor9 days ago

Специалист по технической инфраструктуре, платформа Enterprise Evals

Salary not specified
Полная занятостьОфис

Responsibilities

  • 01Define golden sets: decompose real tasks and encode the expert quality bar.
  • 02Build verifiers over agent trajectories and outputs, calibrated and hard to game.
  • 03Build the eval platform that runs offline environments, task suites, and grading at scale.
  • 04Run loss analysis over production trajectories and turn failure modes into regression tests.
  • 05Run the optimization loop across models, prompts, skills, and harnesses.
  • 06Own the rollout gates that decide whether an agent change ships.
  • 07Partner with the Enterprise Platform team and the Applied AI engineers embedded with customers.

Requirements

  • 01Professional, academic, or research experience in agent engineering and evaluation, including how agent runtimes and harnesses produce a trajectory and where it fails.
  • 02Experience building evaluation suites for LLM or agent systems, and familiarity with how benchmarks such as terminal-bench, tau-bench, and APEX are constructed and where they get gamed.
  • 03Judgment about task and rubric design: turning a fuzzy notion of quality into something measurable, with agent or model improvements to show for it.
  • 04Strong software engineering fundamentals, and the ability to work independently on ambiguous, loosely specified problems.
  • 05Bonus: experience with Harbor environments and RL environments.

What we offer

  • 01Work in-person five days a week in San Francisco, NYC, or London offices.
  • 02Up to $15k relocation bonus.
  • 03$10K housing bonus (if you live within 0.5 miles of office).
  • 04$1.5K monthly stipend for meals.
  • 05Generous equity grant vested over 4 years.
  • 06Free Equinox membership.
  • 07$200 monthly laundry reimbursement.
  • 08$200 monthly personal wellness reimbursement.
  • 09Health, Dental, Vision insurance