Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Research Engineer - Benchmarking

Home > Other programming jobs

Research Engineer - Benchmarking in USA

  • San Francisco, CA

Responsibilities

  • Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning.
  • Build and operate end-to-end LLM evaluation systems, including runs, scoring, dashboards, and reporting.
  • Conduct systematic failure analysis of model outputs, categorize failure modes, quantify their prevalence, and inform reward design, data curation, and benchmark design.
  • Create and refine rubrics, automated evaluators, and scoring frameworks, balancing rigor, scalability, human evaluation, model-as-judge evaluation, calibration, and agreement.
  • Measure data usability, quality, and impact on key benchmarks and guide data generation, augmentation, and curation.
  • Collaborate with AI researchers, applied AI teams, and data producers to align evaluations with training objectives.
  • Own benchmarking, evaluation, and failure-analysis workflows in a fast-paced research environment.

Requirements

  • Strong applied research background focused on model evaluation, benchmarking, and/or failure analysis.
  • Strong coding skills and hands-on experience with ML models and evaluation code.
  • Solid understanding of data structures, algorithms, and backend systems.
  • Comfort working with APIs, SQL/NoSQL, and cloud platforms for running and storing evaluation results.
  • Ability to reason about model behavior, experimental results, and data quality from evaluations and failure analyses.
  • Industry experience on a post-training or evaluation/benchmarking team is preferred.
  • Publications at top-tier venues such as NeurIPS, ICML, or ACL are preferred, especially in evaluation or benchmarking.
  • Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines is preferred.
  • Experience with synthetic data generation, rubric design, or RL-style workflows using evaluations for reward shaping is preferred.
  • Work samples or code demonstrating relevant evaluation frameworks, benchmark suites, failure-analysis reports, or tooling are preferred.

Benefits

  • In-person work five days a week in San Francisco, New York City, or London.
  • Bi-annual performance bonus structure.
  • Generous equity grant vested over four years.
  • Up to $15k relocation bonus.
  • $10K housing bonus for employees living within 0.5 miles of the office.
  • $1.5K monthly meal stipend.
  • Free Equinox membership.
  • $200 monthly laundry reimbursement.
  • $200 monthly personal wellness reimbursement.
  • Health, dental, and vision insurance.

Salary: $130k - $500k/yr

Mercor

We find the best experts in every professional domain and put their knowledge to work training frontier models. Through APEX, we measure whether those models can actually perform economically valuable work. We're also bringing that expertise to enterprises: deploying custom AI agents, staffing te...

Similar positions

Senior Test Automation Engineer

  • City of New York
  • Full time
  • USA
  • 09/13/2026
  • Brooklyn, NY

Backend Engineer, Security Platform

  • Vannevar
  • Full time
  • USA
  • 09/13/2026
  • Salary: $150k - $215k/yr
  • Remote

Agentic Security Engineer

  • NXP Semiconductors
  • Full time
  • USA
  • 09/13/2026
  • Austin, TX

Sr Application Security Engineer

  • Henry Schein One
  • Full time
  • USA
  • 09/13/2026
  • Salary: $125k-$160k
  • Remote

Senior DevOps Engineer, Infrastructure & Reliabili

  • Worth AI
  • Full time
  • USA
  • 09/13/2026
  • Remote