Dataset Viewer
Auto-converted to Parquet Duplicate
query-id
string
call_index
int64
corpus-id
string
C1
int64
C2
int64
C3
int64
C4
int64
C5
int64
finish_reason
string
855410
0
2945339
0
0
0
0
0
stop
855410
0
2164297
0
0
0
0
0
stop
855410
0
7560274
0
0
0
0
0
stop
855410
0
5008542
0
0
0
0
0
stop
855410
0
4151850
0
0
0
0
0
stop
855410
0
7128710
0
0
0
0
0
stop
855410
0
367832
0
0
0
0
0
stop
855410
0
4810235
0
0
0
0
0
stop
855410
0
6425421
0
0
0
0
0
stop
855410
0
1732923
0
0
0
0
0
stop
855410
1
1041939
0
0
0
0
0
stop
855410
1
7238608
0
0
0
0
0
stop
855410
1
6704335
0
0
0
0
0
stop
855410
1
5465332
0
0
0
0
0
stop
855410
1
889438
0
0
0
0
0
stop
855410
1
8373834
0
0
0
0
0
stop
855410
1
7291107
0
0
0
0
0
stop
855410
1
476433
0
0
0
0
0
stop
855410
1
7972234
0
0
0
0
0
stop
855410
1
4263016
0
0
0
0
0
stop
855410
2
983642
0
0
0
0
0
stop
855410
2
681986
0
0
0
0
0
stop
855410
2
8671033
0
0
0
0
0
stop
855410
2
7238783
0
0
0
0
0
stop
855410
2
6707516
0
0
0
0
0
stop
855410
2
7140273
0
0
0
0
0
stop
855410
2
5210914
0
0
0
0
0
stop
855410
2
4299367
0
0
0
0
0
stop
855410
2
7509177
0
0
0
0
0
stop
855410
2
515982
0
0
0
0
0
stop
855410
3
2850382
0
0
0
0
0
stop
855410
3
8651775
1
1
1
1
0
stop
855410
3
4972697
0
0
0
0
0
stop
855410
3
5008545
0
0
0
0
0
stop
855410
3
4833539
0
0
0
0
0
stop
855410
3
3084026
0
0
0
0
0
stop
855410
3
5329993
0
0
0
0
0
stop
855410
3
6306434
0
0
0
0
0
stop
855410
3
5965635
0
0
0
0
0
stop
855410
3
8310088
0
0
0
0
0
stop
855410
4
6734486
0
0
0
0
0
stop
855410
4
6950985
0
0
0
0
0
stop
855410
4
1824280
0
0
0
0
0
stop
855410
4
8671031
0
0
0
0
0
stop
855410
4
8157624
0
0
0
0
0
stop
855410
4
5321528
0
0
0
0
0
stop
855410
4
6481240
0
0
0
0
0
stop
855410
4
3594261
0
0
0
0
0
stop
855410
4
7972231
0
0
0
0
0
stop
855410
4
8352927
0
0
0
0
0
stop
855410
5
4246164
0
0
0
0
0
stop
855410
5
7691815
0
0
0
0
0
stop
855410
5
4420015
0
0
0
0
0
stop
855410
5
5210919
0
0
0
0
0
stop
855410
5
181264
0
0
0
0
0
stop
855410
5
8468783
0
0
0
0
0
stop
855410
5
4103170
0
0
0
0
0
stop
855410
5
4146680
0
0
0
0
0
stop
855410
5
8361990
0
0
0
0
0
stop
855410
5
510943
0
0
0
0
0
stop
855410
6
1494095
0
0
0
0
0
stop
855410
6
7004260
0
0
0
0
0
stop
855410
6
4636144
0
0
0
0
0
stop
855410
6
681990
0
0
0
0
0
stop
855410
6
6947490
0
0
0
0
0
stop
855410
6
4708049
0
0
0
0
0
stop
855410
6
6165926
0
0
0
0
0
stop
855410
6
351153
0
0
0
0
0
stop
855410
6
2531289
0
0
0
0
0
stop
855410
6
3696341
0
0
0
0
0
stop
855410
7
6467175
0
0
0
0
0
stop
855410
7
1956109
0
0
0
0
0
stop
855410
7
7611461
0
0
0
0
0
stop
855410
7
8621844
0
0
0
0
0
stop
855410
7
681984
0
0
0
0
0
stop
855410
7
4151857
0
0
0
0
0
stop
855410
7
8651776
1
0
1
0
0
stop
855410
7
7210883
0
0
0
0
0
stop
855410
7
7031507
0
0
0
0
0
stop
855410
7
984533
0
0
0
0
0
stop
855410
8
197471
0
0
0
0
0
stop
855410
8
5723210
0
0
0
0
0
stop
855410
8
3249944
0
0
0
0
0
stop
855410
8
2147202
0
0
0
0
0
stop
855410
8
3557381
0
0
0
0
0
stop
855410
8
2519616
0
0
0
0
0
stop
855410
8
1075228
0
0
0
0
0
stop
855410
8
8087245
0
0
0
0
0
stop
855410
8
4406031
0
0
0
0
0
stop
855410
8
5381946
0
0
0
0
0
stop
855410
9
5210917
0
0
0
0
0
stop
855410
9
2785222
0
0
0
0
0
stop
855410
9
4635595
0
0
0
0
0
stop
855410
9
681985
0
0
0
0
0
stop
855410
9
1947769
1
1
1
1
0
stop
855410
9
1094091
0
0
0
0
0
stop
855410
9
2519618
0
0
0
0
0
stop
855410
9
8776516
0
0
0
0
0
stop
855410
9
1886069
0
0
0
0
0
stop
855410
9
8407524
0
0
0
0
0
stop
End of preview. Expand in Data Studio

RCP-nDCG · TREC-DL (MTEB-style, with continuous relevance gains)

Evaluate with stock mteb (≥ 2.0.1). No fork is needed. This repo ships rcp_ndcg_tasks.py. It defines one mteb task per subset (2 tasks: TRECDL2019RCPRetrieval, TRECDL2020RCPRetrieval), which reports the continuous-gain metric ndcg_float_at_10 alongside mteb's usual metrics:

import importlib.util
import mteb
from huggingface_hub import hf_hub_download

path = hf_hub_download("fabianschmidt-cohere/rcp-ndcg-trecdl", "rcp_ndcg_tasks.py", repo_type="dataset")
spec = importlib.util.spec_from_file_location("rcp_ndcg_tasks", path)
rcp = importlib.util.module_from_spec(spec)
spec.loader.exec_module(rcp)

tasks = rcp.get_tasks()                        # both subsets; or rcp.get_tasks(["trec_dl_2019"])
model = mteb.get_model("sentence-transformers/all-MiniLM-L6-v2")   # any mteb bi- or cross-encoder
results = mteb.evaluate(model, tasks=tasks)

rcp.get_tasks(mode="retrieval") gives the full-corpus retrieval view instead (bi-encoders only). The same tasks (same names, metadata and numbers) are proposed for mteb itself in PR #5516. If your mteb already ships them, get_tasks() returns the built-in ones.

The two TREC Deep Learning Track passage ranking tasks (2019: 43 queries, 2020: 54 queries) over the MS MARCO passage collection, with the official NIST multi-graded relevance judgments (grades 0–3) and full provenance. Subsets are named trec_dl_2019 / trec_dl_2020, and the split is test, as in mteb's TRECDL2019 / TRECDL2020 tasks.

RCP-nDCG (nDCG over rubric-calibrated preferences) replaces integer relevance labels with calibrated relevance probabilities. An LLM judge compares pooled documents in a listwise tournament and checks each one against a rubric of five binary relevance criteria. A 2PL item-response model then calibrates those verdicts into a latent relevance theta per (query, document), and the relevance probability gain(theta) is the gain in nDCG.

⚠️ Gains come from ONE judge: Qwen3.6-27B-FP8

Every gain value comes from a single LLM judge, Qwen3.6-27B-FP8. Gains come from Qwen3.6-27B-FP8, a different judge from the NanoBEIR/BRIGHT/ViDoRe RCP-nDCG uploads (Qwen3.5-397B-A17B), so RCP scores here are not directly comparable to those suites. The judge answered five binary relevance criteria for each pooled document; a 2PL IRT model was fitted to those answers and calibrated. The gain is the paper's weighted_discrimination link:

gain(θ) = Σ_k γ_k · σ(γ_k (θ − β_k)) / Σ_k γ_k        ∈ (0, 1)
  • item_params for the gain are in provenance/item_params.qwen36-27b.json.
  • Every gain ships with its raw calibrated theta, so a different link can be applied without re-judging.

Recommended use: reranking

{subset}-top_ranked is the judged candidate list, and every candidate has a gain. Loaded through mteb, a model reranks this list. This is the recommended setting:

  • ndcg_at_10: standard nDCG over the upstream NIST qrels (unchanged).
  • ndcg_float_at_10 (the main score): nDCG over the calibrated relevance probabilities, with linear gain. Documents tied on score are credited their tie group's mean gain.

Full-corpus retrieval also works (mode="retrieval"): {subset}-corpus is the complete MS MARCO passage collection (8,841,823 passages). Only the qrels metrics are reported there, because gains exist for the judged pools only.

Evaluation protocol

  • Judged pools. The candidates are the TREC-DL judged pools: the fused first-stage top-150 plus every NIST-judged document (pool sizes 198–585 per query). Every NIST-judged document is in the pool, so there are no unjudged human positives: every pooled document has a gain and a theta, and every human-labelled document has its score.
  • Grades. score is the official NIST relevance grade (0–3), unchanged from the upstream TREC-DL datasets. A score of 0 does not necessarily mean a human judged the document non-relevant: pooled documents outside the NIST judgments also carry score = 0 (filter theta not null to separate judge-scored documents, or score > 0 for the original human qrels).
  • Corpus text. The corpus is the MS MARCO passage collection as in ir_datasets msmarco-passage; its ids are identical to the upstream mteb TREC-DL datasets. That dataset's text carries a double-encoding (mojibake) error in 2,025,452 of 8,841,823 passages (22.9%), which is corrected here: the text in this repo is the text the judge saw.

Layout

config columns notes
{subset}-corpus id, title, text full MS MARCO passage corpus (8,841,823 passages), shared by both subsets
{subset}-queries id, text
{subset}-qrels query-id, corpus-id, score, gain, theta score: NIST grade, unchanged; gain: calibrated relevance probability (see above)
{subset}-top_ranked query-id, corpus-ids judged pool, in pool order
provenance-pool query-id, corpus-id, pool_rank, rrf_score, gt_injected, excluded, judged split = subset; excluded is always false here (no exclusions)
provenance-runs model, is_judge, query-id, corpus-id, score, rank_by_score, stored_position 14 rerankers + 2 judge strategies
provenance-judge-criteria query-id, call_index, corpus-id, C1…C5, finish_reason raw per-call criterion verdicts

Gains live in the qrels. {subset}-qrels has one row per (query, document) in the judged pool:

  • score: the NIST relevance grade (0–3), unchanged; 0 for pooled documents without a NIST label.
  • gain: the calibrated relevance probability, used by ndcg_float_at_k.
  • theta: the raw calibrated latent relevance behind gain.

Stock mteb reads only query-id, corpus-id, score. The extra score-0 rows change none of its retrieval scores (verified on every published run). A caveat: statistics that count "annotated" documents treat the judged pool as annotated, which affects mteb's descriptive statistic unique_relevant_docs and trec_eval's bpref / judged@k.

Query instructions equal the mteb task's prompt["query"] for both originals and are applied by mteb, so no -instruction config is shipped.

How the pool was built

  1. First stage: Cohere Embed v4, BM25 and Octen-Embedding-8B are fused with unweighted RRF (k = 60), keeping the top 150 (provenance-pool.rrf_score).
  2. Judged union: every NIST-judged document was added to the fused top-150 (the TREC-DL judged pools, up to 585 deep on 2019). Documents outside the fused top-150 have rrf_score = 0; qrel positives among them are marked gt_injected (1,844 on 2019, 1,566 on 2020).
  3. Judging: a listwise tournament, then pointwise binary criteria, then a 2PL fit.
    • Stage B rates every pooled document against five binary criteria (C1–C5), and a 2PL model is fitted to these verdicts.
    • The raw verdicts are in provenance-judge-criteria.

Data notes

  • Ties and the paper's numbers. mteb's ndcg_float_at_k gives tied documents their group's mean gain, so the number depends only on the reranker's scores. The paper's TREC-DL tables break ties by the judged pool's order (the fused top 150 first, then the added NIST-judged passages). With that rule, replaying provenance-runs on these gains reproduces the paper's RCP-nDCG@10 and nDCG@10 cells; the two tie rules differ by at most a few thousandths per cell.
  • is_judge = true runs in provenance-runs are the judge itself, not independent systems.
  • The 2PL calibration covers all 597 queries judged in this run (MS MARCO dev, TREC-DL 2019 and 2020) with one shared item-parameter fit; provenance/item_params.qwen36-27b.json holds the parameters used for the gains here.

Licence and attribution

Derived from the TREC-DL tasks in mteb (TRECDL2019, TRECDL2020), which are declared msr-la-nc (Microsoft Research License — non-commercial use only) in their mteb task metadata. Corpus, queries and qrels are redistributed unchanged at the pinned upstream revisions (the corpus text carries the encoding correction described above). The provenance-runs scores come from third-party reranker models and APIs, and the gains from an LLM judge's outputs, so those providers' terms also apply. Check the upstream terms before commercial use. Please cite MTEB and MS MARCO / TREC-DL alongside RCP-nDCG (arXiv:2609.35739):

@misc{schmidt2026rcp,
  title         = {Rubric-Calibrated Preferences: Cross-Query Calibration of LLM Judgments via Item Response Theory},
  author        = {Schmidt, Fabian David and Crisostomi, Donato and Lassance, Carlos and Reimers, Nils},
  year          = {2026},
  eprint        = {2609.35739},
  archiveprefix = {arXiv},
  primaryclass  = {cs.IR},
  url           = {https://arxiv.org/abs/2609.35739},
}

The blind human study and the external LLM-judge comparisons that validate RCP-nDCG are released in fabianschmidt-cohere/rcp-ndcg-external-validation.

Downloads last month
648

Paper for fabianschmidt-cohere/rcp-ndcg-trecdl