Mercor6 days ago
Research Engineer – Benchmarking
Salary not specified
MARKET
30,208 ₽median for this role
AI Researcher · 30 jobs with disclosed pay
5,600half of the offers: 16,410–78,975131,300
The employer didn't disclose pay — compare with the market yourself.
Полная занятостьОфис
Responsibilities
- 01Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning; ensure benchmarks scale with training and stay aligned with product and research goals
- 02Build and operate LLM evaluation systems end-to-end runs, scoring, dashboards, and reporting, so researchers and applied AI teams can track model performance and compare runs at scale
- 03Run systematic failure analysis on model outputs (e.g., wrong tool use, reasoning errors, safety/alignment issues); categorize failure modes, quantify prevalence, and feed findings into reward design, data curation, and benchmark design
- 04Create and refine rubrics, automated evaluators, and scoring frameworks that drive training and evaluation decisions; balance rigor with scalability (human vs. model-as-judge, calibration, agreement)
- 05Quantify data usability, quality, and impact on key benchmarks; use evals and failure analysis to guide data generation, augmentation, and curation
- 06Work with AI researchers, applied AI teams, and data producers to align evals with training objectives and to prioritize benchmarks and failure analyses that matter most
- 07Operate in a high-iteration research setting with strong ownership of benchmarks, evals, and failure-analysis workflows
Requirements
- 01Strong applied research background, with focus on model evaluation, benchmarking, and/or failure analysis
- 02Strong coding skills and hands-on experience with ML models and evaluation code
- 03Solid grasp of data structures, algorithms, and backend systems
- 04Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing eval results
- 05Ability to reason about model behavior, experimental results, and data quality from evals and failure analyses
- 06Excitement to work in person in San Francisco five days a week in a high-intensity, high-ownership environment
- 07Industry experience on a post-training or evaluation/benchmarking team (highest priority)
- 08Publications at top-tier venues (NeurIPS, ICML, ACL), especially in evaluation or benchmarking
- 09Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines
- 10Experience with synthetic data generation, rubric design, or RL-style workflows that use evals for reward shaping
- 11Work samples or code (e.g., eval frameworks, benchmark suites, failure-analysis reports or tooling) that demonstrate relevant skills
What we offer
- 01Bi-annual performance bonus structure
- 02Generous equity grant vested over 4 years
- 03Up to $15k Relocation bonus
- 04$10K housing bonus (if you live within 0.5 miles of our office)
- 05$1.5K monthly stipend for meals
- 06Free Equinox membership
- 07$200 monthly laundry reimbursement
- 08$200 monthly personal wellness reimbursement
- 09Health, Dental, Vision insurance
- 10In-person work five days a week in San Francisco, NYC, or London offices