Mercor9 days ago
Специалист по технической инфраструктуре, платформа Enterprise Evals
Salary not specified
Полная занятостьОфис
Responsibilities
- 01Define golden sets: decompose real tasks and encode the expert quality bar.
- 02Build verifiers over agent trajectories and outputs, calibrated and hard to game.
- 03Build the eval platform that runs offline environments, task suites, and grading at scale.
- 04Run loss analysis over production trajectories and turn failure modes into regression tests.
- 05Run the optimization loop across models, prompts, skills, and harnesses.
- 06Own the rollout gates that decide whether an agent change ships.
- 07Partner with the Enterprise Platform team and the Applied AI engineers embedded with customers.
Requirements
- 01Professional, academic, or research experience in agent engineering and evaluation, including how agent runtimes and harnesses produce a trajectory and where it fails.
- 02Experience building evaluation suites for LLM or agent systems, and familiarity with how benchmarks such as terminal-bench, tau-bench, and APEX are constructed and where they get gamed.
- 03Judgment about task and rubric design: turning a fuzzy notion of quality into something measurable, with agent or model improvements to show for it.
- 04Strong software engineering fundamentals, and the ability to work independently on ambiguous, loosely specified problems.
- 05Bonus: experience with Harbor environments and RL environments.
What we offer
- 01Work in-person five days a week in San Francisco, NYC, or London offices.
- 02Up to $15k relocation bonus.
- 03$10K housing bonus (if you live within 0.5 miles of office).
- 04$1.5K monthly stipend for meals.
- 05Generous equity grant vested over 4 years.
- 06Free Equinox membership.
- 07$200 monthly laundry reimbursement.
- 08$200 monthly personal wellness reimbursement.
- 09Health, Dental, Vision insurance