Datasets:
query-id string | call_index int64 | corpus-id string | C1 int64 | C2 int64 | C3 int64 | C4 int64 | C5 int64 | finish_reason string |
|---|---|---|---|---|---|---|---|---|
855410 | 0 | 2945339 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 0 | 2164297 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 0 | 7560274 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 0 | 5008542 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 0 | 4151850 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 0 | 7128710 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 0 | 367832 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 0 | 4810235 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 0 | 6425421 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 0 | 1732923 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 1041939 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 7238608 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 6704335 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 5465332 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 889438 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 8373834 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 7291107 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 476433 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 7972234 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 1 | 4263016 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 983642 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 681986 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 8671033 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 7238783 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 6707516 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 7140273 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 5210914 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 4299367 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 7509177 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 2 | 515982 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 3 | 2850382 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 3 | 8651775 | 1 | 1 | 1 | 1 | 0 | stop |
855410 | 3 | 4972697 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 3 | 5008545 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 3 | 4833539 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 3 | 3084026 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 3 | 5329993 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 3 | 6306434 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 3 | 5965635 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 3 | 8310088 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 6734486 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 6950985 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 1824280 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 8671031 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 8157624 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 5321528 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 6481240 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 3594261 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 7972231 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 4 | 8352927 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 4246164 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 7691815 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 4420015 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 5210919 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 181264 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 8468783 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 4103170 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 4146680 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 8361990 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 5 | 510943 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 1494095 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 7004260 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 4636144 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 681990 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 6947490 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 4708049 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 6165926 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 351153 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 2531289 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 6 | 3696341 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 7 | 6467175 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 7 | 1956109 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 7 | 7611461 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 7 | 8621844 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 7 | 681984 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 7 | 4151857 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 7 | 8651776 | 1 | 0 | 1 | 0 | 0 | stop |
855410 | 7 | 7210883 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 7 | 7031507 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 7 | 984533 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 197471 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 5723210 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 3249944 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 2147202 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 3557381 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 2519616 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 1075228 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 8087245 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 4406031 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 8 | 5381946 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 9 | 5210917 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 9 | 2785222 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 9 | 4635595 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 9 | 681985 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 9 | 1947769 | 1 | 1 | 1 | 1 | 0 | stop |
855410 | 9 | 1094091 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 9 | 2519618 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 9 | 8776516 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 9 | 1886069 | 0 | 0 | 0 | 0 | 0 | stop |
855410 | 9 | 8407524 | 0 | 0 | 0 | 0 | 0 | stop |
RCP-nDCG · TREC-DL (MTEB-style, with continuous relevance gains)
Evaluate with stock mteb (≥ 2.0.1). No fork is
needed. This repo ships rcp_ndcg_tasks.py. It defines one mteb task per subset (2 tasks:
TRECDL2019RCPRetrieval, TRECDL2020RCPRetrieval), which reports the continuous-gain metric
ndcg_float_at_10 alongside mteb's usual metrics:
import importlib.util
import mteb
from huggingface_hub import hf_hub_download
path = hf_hub_download("fabianschmidt-cohere/rcp-ndcg-trecdl", "rcp_ndcg_tasks.py", repo_type="dataset")
spec = importlib.util.spec_from_file_location("rcp_ndcg_tasks", path)
rcp = importlib.util.module_from_spec(spec)
spec.loader.exec_module(rcp)
tasks = rcp.get_tasks() # both subsets; or rcp.get_tasks(["trec_dl_2019"])
model = mteb.get_model("sentence-transformers/all-MiniLM-L6-v2") # any mteb bi- or cross-encoder
results = mteb.evaluate(model, tasks=tasks)
rcp.get_tasks(mode="retrieval") gives the full-corpus retrieval view instead (bi-encoders only).
The same tasks (same names, metadata and numbers) are proposed for mteb itself in
PR #5516. If your mteb already ships them,
get_tasks() returns the built-in ones.
The two TREC Deep Learning Track passage ranking tasks (2019: 43 queries, 2020: 54 queries) over the
MS MARCO passage collection, with the official NIST multi-graded relevance judgments (grades 0–3) and
full provenance. Subsets are named trec_dl_2019 / trec_dl_2020, and the split is test, as in
mteb's TRECDL2019 / TRECDL2020 tasks.
RCP-nDCG (nDCG over rubric-calibrated preferences) replaces integer relevance labels with
calibrated relevance probabilities. An LLM judge compares pooled documents in a listwise tournament and
checks each one against a rubric of five binary relevance criteria. A 2PL item-response model then
calibrates those verdicts into a latent relevance theta per (query, document), and the relevance
probability gain(theta) is the gain in nDCG.
⚠️ Gains come from ONE judge: Qwen3.6-27B-FP8
Every gain value comes from a single LLM judge, Qwen3.6-27B-FP8. Gains come from Qwen3.6-27B-FP8,
a different judge from the NanoBEIR/BRIGHT/ViDoRe RCP-nDCG uploads (Qwen3.5-397B-A17B), so RCP scores
here are not directly comparable to those suites. The judge answered five binary relevance criteria for
each pooled document; a 2PL IRT model was fitted to those answers and calibrated. The gain is the
paper's weighted_discrimination link:
gain(θ) = Σ_k γ_k · σ(γ_k (θ − β_k)) / Σ_k γ_k ∈ (0, 1)
item_paramsfor the gain are inprovenance/item_params.qwen36-27b.json.- Every
gainships with its raw calibratedtheta, so a different link can be applied without re-judging.
Recommended use: reranking
{subset}-top_ranked is the judged candidate list, and every candidate has a gain. Loaded through
mteb, a model reranks this list. This is the recommended setting:
ndcg_at_10: standard nDCG over the upstream NIST qrels (unchanged).ndcg_float_at_10(the main score): nDCG over the calibrated relevance probabilities, with linear gain. Documents tied on score are credited their tie group's mean gain.
Full-corpus retrieval also works (mode="retrieval"): {subset}-corpus is the complete MS MARCO
passage collection (8,841,823 passages). Only the qrels metrics are reported there, because gains
exist for the judged pools only.
Evaluation protocol
- Judged pools. The candidates are the TREC-DL judged pools: the fused first-stage top-150 plus
every NIST-judged document (pool sizes 198–585 per query). Every NIST-judged document is in the
pool, so there are no unjudged human positives: every pooled document has a
gainand atheta, and every human-labelled document has itsscore. - Grades.
scoreis the official NIST relevance grade (0–3), unchanged from the upstream TREC-DL datasets. Ascoreof 0 does not necessarily mean a human judged the document non-relevant: pooled documents outside the NIST judgments also carryscore = 0(filterthetanot null to separate judge-scored documents, orscore > 0for the original human qrels). - Corpus text. The corpus is the MS MARCO passage collection as in ir_datasets
msmarco-passage; its ids are identical to the upstream mteb TREC-DL datasets. That dataset's text carries a double-encoding (mojibake) error in 2,025,452 of 8,841,823 passages (22.9%), which is corrected here: the text in this repo is the text the judge saw.
Layout
| config | columns | notes |
|---|---|---|
{subset}-corpus |
id, title, text |
full MS MARCO passage corpus (8,841,823 passages), shared by both subsets |
{subset}-queries |
id, text |
|
{subset}-qrels |
query-id, corpus-id, score, gain, theta |
score: NIST grade, unchanged; gain: calibrated relevance probability (see above) |
{subset}-top_ranked |
query-id, corpus-ids |
judged pool, in pool order |
provenance-pool |
query-id, corpus-id, pool_rank, rrf_score, gt_injected, excluded, judged |
split = subset; excluded is always false here (no exclusions) |
provenance-runs |
model, is_judge, query-id, corpus-id, score, rank_by_score, stored_position |
14 rerankers + 2 judge strategies |
provenance-judge-criteria |
query-id, call_index, corpus-id, C1…C5, finish_reason |
raw per-call criterion verdicts |
Gains live in the qrels. {subset}-qrels has one row per (query, document) in the judged pool:
score: the NIST relevance grade (0–3), unchanged; 0 for pooled documents without a NIST label.gain: the calibrated relevance probability, used byndcg_float_at_k.theta: the raw calibrated latent relevance behindgain.
Stock mteb reads only query-id, corpus-id, score. The extra score-0 rows change none of its
retrieval scores (verified on every published run). A caveat: statistics that count "annotated"
documents treat the judged pool as annotated, which affects mteb's descriptive statistic
unique_relevant_docs and trec_eval's bpref / judged@k.
Query instructions equal the mteb task's prompt["query"] for both originals and are applied by mteb,
so no -instruction config is shipped.
How the pool was built
- First stage: Cohere Embed v4, BM25 and
Octen-Embedding-8B are fused with unweighted RRF
(
k = 60), keeping the top 150 (provenance-pool.rrf_score). - Judged union: every NIST-judged document was added to the fused top-150 (the TREC-DL judged
pools, up to 585 deep on 2019). Documents outside the fused top-150 have
rrf_score = 0; qrel positives among them are markedgt_injected(1,844 on 2019, 1,566 on 2020). - Judging: a listwise tournament, then pointwise binary criteria, then a 2PL fit.
- Stage B rates every pooled document against five binary criteria (C1–C5), and a 2PL model is fitted to these verdicts.
- The raw verdicts are in
provenance-judge-criteria.
Data notes
- Ties and the paper's numbers. mteb's
ndcg_float_at_kgives tied documents their group's mean gain, so the number depends only on the reranker's scores. The paper's TREC-DL tables break ties by the judged pool's order (the fused top 150 first, then the added NIST-judged passages). With that rule, replayingprovenance-runson these gains reproduces the paper's RCP-nDCG@10 and nDCG@10 cells; the two tie rules differ by at most a few thousandths per cell. is_judge = trueruns inprovenance-runsare the judge itself, not independent systems.- The 2PL calibration covers all 597 queries judged in this run (MS MARCO dev, TREC-DL 2019 and 2020)
with one shared item-parameter fit;
provenance/item_params.qwen36-27b.jsonholds the parameters used for the gains here.
Licence and attribution
Derived from the TREC-DL tasks in mteb (TRECDL2019, TRECDL2020), which are declared
msr-la-nc (Microsoft Research License — non-commercial use only) in their mteb task metadata.
Corpus, queries and qrels are redistributed unchanged at the pinned upstream revisions (the corpus
text carries the encoding correction described above). The provenance-runs scores come from
third-party reranker models and APIs, and the gains from an LLM judge's outputs, so those providers'
terms also apply. Check the upstream terms before commercial use. Please cite MTEB and MS MARCO /
TREC-DL alongside RCP-nDCG (arXiv:2609.35739):
@misc{schmidt2026rcp,
title = {Rubric-Calibrated Preferences: Cross-Query Calibration of LLM Judgments via Item Response Theory},
author = {Schmidt, Fabian David and Crisostomi, Donato and Lassance, Carlos and Reimers, Nils},
year = {2026},
eprint = {2609.35739},
archiveprefix = {arXiv},
primaryclass = {cs.IR},
url = {https://arxiv.org/abs/2609.35739},
}
The blind human study and the external LLM-judge comparisons that validate RCP-nDCG are released in
fabianschmidt-cohere/rcp-ndcg-external-validation.
- Downloads last month
- 648