Evaluation

Turning Real User Traffic Into an Eval Set That Catches Your Regressions

OpenRouter's five-step tutorial turns production traffic into a versioned regression suite for prompts and model swaps.

Turning Real User Traffic Into an Eval Set That Catches Your Regressions — article cover

Every prompt edit has the same risk profile as a schema change: it looks fine until something downstream breaks. OpenRouter’s tutorial from September 30, 2026, Building a Golden Eval Dataset from Production Traffic, lays out a concrete process for catching those breaks before deploy — using the traffic you already have.

Public benchmarks won’t catch your regressions

The problem OpenRouter starts from is familiar: you change a prompt, or a provider ships a new checkpoint under the same model ID, and quality drifts in ways MMLU will never measure. A benchmark tests general capability. A golden eval set tests whether your product still behaves.

The definition is worth internalizing: a curated set of production inputs paired with reviewed expected outputs, versioned in Git, and run before every meaningful change. OpenRouter frames it as a regression test suite for behavior rather than code — the difference from a unit test is that the expected output is a judgment call, so your harness has to grade, not just diff strings.

This is the same mindset I argued for in Treat Every Prompt Edit Like a Schema Migration: changes to model behavior deserve the same gate discipline as changes to data shape.

Why production beats synthetic

OpenRouter gives three reasons the foundation should be real traffic:

  • The distribution is right. If 70% of traffic is about pricing, 70% of your eval set should be too. That’s what makes a regression score reflect user impact.
  • The failure modes are real. Users paste error logs, mix three languages in one message, reference a feature you renamed. Synthetic sets only contain failures someone imagined.
  • Examples track your product. A set drawn from live traffic drifts with your product; one frozen at launch doesn’t.

Synthetic data still earns a spot — as a minority filler for rare failure modes you lack real examples of, marked synthetic: true in metadata.

The five steps, compressed

  1. Sample production traffic. One to two weeks of logged inputs, with every field you might slice on later (intent, model version, prompt version) logged at capture time — you can’t reconstruct metadata afterward. Scrub PII before it reaches the harness.
  2. Deduplicate and cluster. Exact match misses paraphrases, so compare embeddings. Keep one of fifty password-reset questions, not all of them, and weight toward failure modes you can’t afford.
  3. Add expected outputs. Usually a rubric, not a string. Use binary criteria (MET/UNMET) and analytic rubrics that score each criterion separately, so a failure tells you what got worse. Have two people annotate a subset; disagreements mean the rubric is broken, not the annotators.
  4. Run your current model and prune. Sort failures into real failures (keep), rubric problems (fix the expected output), and ambiguous items (cut). Expect to cut a chunk — that cut is the point. A set that produces the same pass rate no matter what you change measures nothing.
  5. Commit to Git and wire into CI. Fail the build on a pass-rate drop beyond a threshold; require written acknowledgment on smaller drops. Refresh on a cadence — quarterly review of items older than six months is a reasonable default.

On sizing, OpenRouter cites Langfuse’s guidance: roughly 10 items to explore a single issue, 100 to 1,000 for CI checks on larger changes. Keep a small fast subset as the PR gate and reserve the full set for release branches or nightly runs.

Two details that separate good eval sets from decoration

First, version the dataset, the rubric, and the baseline together. If someone edits a rubric criterion to silence a false regression, that edit should show up in a PR diff — not as an unexplained pass-rate move six weeks later.

Second, don’t retire items just because they keep passing. A passing item is proof the behavior still holds. Retire only when the behavior no longer exists or the expected output is now wrong.

The tutorial closes with cross-model benchmarking through OpenRouter’s API — running the same golden set against candidate models so you pick your next model on evidence from your own traffic instead of leaderboard rank. The full post includes a concrete dataset.jsonl schema and directory layout; the supplied RSS excerpt cuts off before that section’s detail, so read the original for those specifics. If you ship prompt changes regularly, though, the takeaway stands on its own: the fastest path to trustworthy evals runs through the traffic you’re already logging.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

Found this useful?

Support more practical AI articles, tutorials, and build notes.

Buy us a coffee
SHAREXEMAIL