AGENTIC COMMONSAI industry briefings

繁中EN

TOPICEnterprise AIPUBLISHED 2026-10-09

All English articlesInformation Retrieval

Your Retrieval Benchmark May Be Judging Yesterday's Results

On this page6 sections
  1. The problem is the answer key, not the metric
  2. What the human study found
  3. How RCP-nDCG@10 works
  4. Does it match what humans prefer?
  5. What this means for your roadmap
  6. Sources

If you build search or RAG products, you probably trust nDCG@10 as your ground truth. A post published by Cohere on September 30, 2026 argues that trust is increasingly misplaced — and they have human study data to back it up.

The problem is the answer key, not the metric

nDCG rewards putting relevant documents near the top, which is why it became the standard on benchmarks like MTEB and BEIR. But it doesn’t judge relevance itself. It compares your model’s output against qrels — relevance judgments produced by human assessors who reviewed a subset of documents surfaced by earlier retrieval systems.

That pooling was a reasonable compromise when judging every query-document pair was prohibitively expensive. The trade-off is a coverage gap: a newer model can retrieve a genuinely relevant document that was never judged, and conventional nDCG scores it as irrelevant. Cohere illustrates this with a ViDoRe v3 query about nitrogen tank color codes — the document at rank 3 quotes the answer directly, but because it was never labeled, the nDCG@5 comes out at 0.760 based on just one judged result.

In short: as your retrieval gets better at exploring the corpus, the benchmark increasingly measures what previous generations could find, not what your system actually returns.

What the human study found

Cohere ran a 46-person annotation study to test how well existing labels predict real human judgment. The numbers are uncomfortable reading for anyone relying on legacy benchmarks:

  • Qrels achieved an AUC of only 0.65 as a predictor of blind human relevance grades — barely better than chance for something treated as ground truth.
  • Among documents labeled irrelevant, 28% were judged useful by humans. Missed relevance, not false positives, is the dominant error.
  • Sparse labeling is the driver: benchmarks tagged 29% of documents as relevant; human evaluators said 60%.
  • Relevance itself is subjective — independent reviewers agreed exactly on a 0-4 scale only 42% of the time.

How RCP-nDCG@10 works

The proposed fix, Rubric-Calibrated Preferences nDCG@10, replaces the fixed answer key with a calibrated AI judge. Instead of checking results against what was previously labeled, it scores every retrieved document on its actual relevance against an explicit rubric.

Two things make this more than asking an LLM to grade documents. The judge answers five yes/no rubric questions per document — the same questions for every query — and also compares documents against each other in groups. Calibration combines the signals: comparisons set the order, the rubric places it on a scale shared across queries. The result is a relevance score that means the same thing regardless of query, which a raw LLM grade can’t provide.

Does it match what humans prefer?

In a blind study across NanoBEIR, BRIGHT, and ViDoRe v3 — 289 head-to-head contests over 273 queries, with annotators unaware of system names or answer keys — the results favored the new approach:

  • Where the two metrics disagreed, reviewers sided with RCP-nDCG 70% of the time versus 52% for conventional nDCG (in a pool that deliberately over-samples disagreements).
  • Agreement rises with the reported margin, hitting 97% at the widest gaps. Below a margin of about 0.02, it’s a coin toss.
  • The biggest win comes from crediting relevant documents the answer key missed — worth 19 percentage points, with rubric grading adding 6 more.

Kenneth Enevoldsen, primary maintainer of MTEB, is quoted in the post supporting the approach and bringing it into MTEB.

What this means for your roadmap

Cohere says it optimized its upcoming fifth-generation Embed and Rerank against RCP-nDCG@10, and cautions those models may not top legacy nDCG leaderboards as a result. The paper, code, and data are public, so you can evaluate the methodology rather than take the framing on faith.

The practical takeaway: treat leaderboard deltas between top-tier retrievers with suspicion, since the labels may lack the resolution to distinguish them. If you’re shipping a RAG product, benchmarking on your own corpus with human spot-checks still beats any public metric — including a new one. This also connects to a theme I’ve covered before: agent benchmarks are only as trustworthy as the search they actually depend on (ATLAS).

One caveat worth keeping in mind: an AI judge introduces its own biases, and the study’s contests over-sample disagreements between metrics. Validation against human annotators is encouraging, but your domain — especially specialized enterprise content — may behave differently. The code and data are on GitHub; running it on a slice of your own queries is a reasonable next step before betting a model choice on it.

Sources

AGENTIC COMMONSOperated by PHLEGON LABS
SHAREXEMAIL
Support us

Related reading

  1. Generative AI for Business: A Practical Guide to Adoption

    Learn how to adopt generative AI in business: key use cases, benefits, challenges, and a three-step framework for successful implementation.

    Enterprise AI

  2. What Anthropic's $150M Genesis Mission Pledge Changes for AI-Assisted Science

    Anthropic will spend $150M over three years putting Claude, Claude Code, and API credits into federal research projects.

    Anthropic

  3. A Trillion-Parameter Open-Weight Bet: What Mistral's ML4 Preview Means for Builders

    Mistral's 1T-parameter ML4 preview pairs frontier cyber and coding benchmarks with open weights you can self-host by month's end.

    Mistral