AI Agents

Treat Every Prompt Edit Like a Schema Migration

OpenRouter's guide argues agents need locked case sets, structural assertions, and pinned model slugs to catch regressions after prompt or model changes.

Treat Every Prompt Edit Like a Schema Migration — article cover

The scariest changes to an AI agent are the ones your test suite has no reason to look at. You reword one sentence in a system prompt, swap in a new chunking strategy, or bump the model behind a latest alias — and nothing errors out. The transcript still reads fluently. But your support agent now paraphrases the refund policy from memory instead of quoting it, because the paragraph it relied on fell outside the retrieved chunk.

That gap is what OpenRouter’s tutorial from September 30, 2026 addresses. The guide on agent regression testing lays out a workflow for re-running a locked set of cases whenever a prompt, model, tool definition, or retrieval setting changes — then checking results against a written behavioral contract rather than a text diff.

Why code-style regression tests break on agents

Code regression testing assumes a known input, a known correct output, and a diff that flags deviations. Agents break all three assumptions, per OpenRouter:

  • Two correct answers rarely look alike, so diffing against a golden string fails on behavior that was never wrong.
  • The model itself is a moving part. An alias like ~author/family-latest can resolve to a new version without any commit in your repository.
  • A pass expires the moment the baseline moves.

The fix is to assert on structure instead of prose: did the agent call the right tool with the right arguments, respect policy, and ask for missing information? OpenRouter also points out that the model field in every OpenRouter response tells you the concrete model that actually served the request — the cheapest way to notice the thing answering your calls is no longer the thing you tested.

Lock the cases, write the contract

The case set comes first. Cover your common requests, a few edge cases (ambiguous input, policy-boundary requests), and at least one case built to test an invariant you never want broken — a routine refund, a request with no order ID, and a refund above the limit, for a support agent.

Then stop editing the set casually. OpenRouter’s framing is the part worth stealing: every reword turns the set into a new experiment, so treat case changes the way you’d treat a schema migration.

Each case gets two assertions. A structural assertion says what the agent should do — call lookup_order first, leave escalate_to_human alone on a routine refund. A hard invariant says what it must never do — approve a refund above your policy limit without a human. The hard invariant is your automatic ship-blocker, with no threshold and no judgment call attached. For open-ended answers where structure won’t suffice, the guide mentions their Ori Eval tooling supports LLM-judged checks alongside tool-call assertions like run.tool('escalate_to_human').toBeCalled().

// hard invariant as a plain assertion
const called = (data.choices[0].message.tool_calls ?? []).map(
  (c: { function: { name: string } }) => c.function.name,
);
if (!called.includes("escalate_to_human")) {
  throw new Error("hard invariant broken: refund above the limit");
}

The model-swap case: hold everything but one variable

When comparing models, the discipline is simple and easy to skip: keep the prompt, tools, cases, judge, and inference parameters identical, and vary only the model. Use a concrete slug such as anthropic/claude-fable-5.1, never an alias — otherwise you’re comparing against a moving target.

Then read the baseline column before the candidate column. A case that fails on both sides means your test is broken, not your agent. A hard invariant that breaks only on the candidate stops the release. This is the same discipline I’d apply to any model-routing decision — it’s the difference between a defensible choice and vibes, and it pairs well with reading cost curves rather than benchmarks when picking between options, as covered in Sonnet 5.5’s cost-tiering story.

Don’t wire evals into unit tests

One operational detail that matters: OpenRouter’s Ori Eval docs caution that evals hit real models and cost money, so run them in a separate job — triggered manually, on a schedule, or scoped to the paths holding your prompts and model config. ori eval --pilot 1 samples one case per file and reports measured cost per model, so you can price the suite before committing. A failed eval returns a non-zero exit code, which is what lets a worse agent actually block a release.

The practical takeaway: if you ship agents, your regression suite is really a contract test suite with pinned model slugs and locked cases. Start with three to five cases and one hard invariant — that alone will catch most of the silent drift the next prompt edit introduces.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

Found this useful?

Support more practical AI articles, tutorials, and build notes.

Buy us a coffee
SHAREXEMAIL