RuVerBench
August 28, 2026 ยท View on GitHub
RuVerBench is a rubric verification benchmark for evaluating LLM-as-a-Judge reliability in long agentic scenarios.
The benchmark asks a judge model to decide whether an output satisfies an individual rubric. Each judgment is evaluated against a human gold label.
Code and data are licensed separately. See LICENSE for code and
DATA_LICENSE.md plus THIRD_PARTY_NOTICES.md for dataset terms and
third-party notices. The data package contains materials from multiple upstream
sources and is not covered by a single blanket license.
Domains
- DeepResearch: question-report pairs.
- AgenticCoding: serialized agent trajectories. The trajectory file contains 217 records; the scored dataset selects 210 cases.
The scored benchmark contains 494 cases and 2,458 rubric-verification instances:
- DeepResearch: 284 scored cases, 1,615 rubric points.
- AgenticCoding: 210 cases, 843 checklist items.
The packaged DeepResearch source files contain 298 question-report records. The main benchmark and leaderboard use the 284 records that have final rubric taxonomy assignments.
Released benchmark files contain normalized inputs, rubrics/checklists, labels, responses/trajectories, and taxonomy assignments. The fixed strategy subset uses the same compact label schema as the main benchmark.
DeepResearch records use the prefixes DRB2_*, RB_*, and RR_* to identify
DeepResearch-Bench-II, ResearcherBench, and ResearchRubrics sources,
respectively. AgenticCoding records are derived from OctoBench. See
DATA_LICENSE.md for the applicable terms.
Repository Layout
RuVerBench/
README.md
RUN_STEPS.md
DATASET_CARD.md
DATA_LICENSE.md
THIRD_PARTY_NOTICES.md
requirements.txt
data/
benchmark/ benchmark inputs, labels, and taxonomy files
predictions/examples/ example judge-prediction files
strategy_fixed20_subset/ fixed subset files for strategy analyses
code/
main_leaderboard/ leaderboard recomputation
run_judges/ API runtime and domain judge runners
strategies/ strategy generation and aggregation scripts
figures/ paper-result figure generation
validate_package.py packaged-data and artifact validation
results/
main_leaderboard/ recomputed leaderboard tables
dataset/ dataset/category summaries
strategies/ prompt, batch, and voting summaries
error_analysis/ model-error overlap files
figures/ generated paper-result figures
tests/
test_judge_runners.py schema and parser tests
Local Reproduction
Install dependencies:
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install --upgrade pip
python3 -m pip install -r requirements.txt
Regenerate figures from exported result tables:
python3 code/figures/plot_paper_figures.py
This command uses exported result tables and does not call external model APIs.
Run the local runner tests without an API key:
python3 -m unittest discover -s tests -v
python3 code/validate_package.py
The repository includes exported leaderboard tables under
results/main_leaderboard/. The complete prediction files for the 18 paper
models are not packaged; example schemas are provided under
data/predictions/examples/. Generate a new prediction file with a judge
runner and score it with code/main_leaderboard/compute_main_leaderboard.py.
Strategy tables are exported under results/strategies/. The fixed subset used
by the strategy analyses is packaged under data/strategy_fixed20_subset/.
Fresh API-based strategy runs are supported by the scripts in code/strategies/,
but exact raw regeneration can change with endpoint behavior, model version,
prompt settings, sampling settings, and run date.
API-Based Regeneration
Fresh judge predictions require an OpenAI-compatible API endpoint:
export JUDGE_MODEL="your-model-id"
export JUDGE_BASE_URL="https://your-endpoint/v1"
export JUDGE_API_KEY="your-api-key"
Then run:
bash code/run_judges/run_deepresearch_judge.sh
bash code/run_judges/run_agenticcoding_judge.sh
Score one generated prediction file:
python3 code/main_leaderboard/compute_main_leaderboard.py \
--domain deepresearch \
--prediction-file outputs/generated_predictions/deepresearch/rubric_eval/<judge-model>/deepresearch_responses_evaluation_results.json
Reproducibility Boundary
This release supports:
- inspecting the exported main leaderboard tables;
- inspecting exported paper strategy tables;
- inspecting the fixed subset files used by the strategy analyses;
- regenerating result figures from packaged tables;
- generating and scoring new judge predictions;
- rerunning prompt, batching, and voting protocols on the fixed subset.
Recomputing the complete paper leaderboard requires the full prediction files for all reported models. Those model outputs are not included in this release.
Fresh model generation is outside the deterministic reproduction path. Results from new API calls depend on endpoint behavior, model version, prompt settings, sampling settings, and run date.
Citation
If you use RuVerBench, please cite the RuVerBench paper and repository. Please
also cite the upstream benchmark papers corresponding to the records used in
your experiment. See THIRD_PARTY_NOTICES.md for the source mapping and
citation links.
Main Files To Read
DATASET_CARD.mdRUN_STEPS.mdresults/main_leaderboard/main_leaderboard.mdresults/strategies/results/error_analysis/