The tradeoff this removes
Embedding-heavy retrieval usually forces a choice: run your best model on everything, or accept lower quality on the query path to keep latency and cost down. Indexing happens once, but every search pays embedding cost again — and agentic workflows multiply that, since a single task can trigger dozens of retrieval calls.
Cohere’s Embed 5, released September 30, 2026, attacks that tradeoff directly. The family has two tiers — Pro at $0.12 per million tokens for text, Fast at $0.08 — and the interesting part isn’t the pricing split. It’s that per Cohere, both tiers share a single embedding space, so vectors from either model compare directly. You can index with Pro and query with Fast without rebuilding the index, or flip the pattern for the opposite tradeoff.
What the shared space is worth in practice
The cross-model claim is easy to make and easy to fake, so the numbers matter. Cohere reports testing every corpus/query pairing across 40 development datasets spanning text, image, fused, and parsed-document retrieval. Mixed combinations stayed close to same-model baselines: 1.6% average loss with Fast queries against a Pro index, 2.7% the other way, with no dataset showing a major failure.
That translates into a deployment pattern Cohere itself recommends for many customers: index with Pro, query with Fast. You keep most of the quality gain of an all-Pro system while making the per-request cost the cheap one. For teams running high-volume RAG or agent loops, that’s a meaningful lever on both latency and spend.
Where the quality claims actually come from
Most of the benchmark story centers on messy enterprise documents. Embed 5 Pro averages 85.8 on ViDoRe V3, which covers visually rich material like financial filings, technical manuals, and regulatory reports — an 8.8-point gain over Embed 4, ahead of Voyage 4 Large (83.7), Gemini Embedding 2 (83.2), and OpenAI text-embedding-3-large (75.5). On financial benchmarks specifically (FinanceBench, FinQA, ViDoRe V3 Finance), Pro ranks first and Fast second on all three.
The parsed-PDF results are worth a closer look if your pipeline converts documents to text before embedding. Cohere argues that parsing strips structure — tables lose row-column relationships, multi-column layouts scramble reading order, charts vanish — and its parsed-document suite reflects that harder problem. Pro averages 84.8 there, with Fast at 83.4, also ahead of Gemini Embedding 2.
One caveat worth flagging: Embed 5 is the first family evaluated with RCP-nDCG@10, a new methodology that scores retrieved documents against query-specific relevance criteria instead of fixed labels. I’ve written before about why your retrieval benchmark may be judging yesterday’s results — and this is exactly the situation where a vendor’s headline metric comes from its own brand-new yardstick. The comparisons against other models on it are Cohere’s own; treat them as a strong signal, not independent validation.
The rest of the spec, briefly
Both tiers take text, images, and fused text-plus-image inputs, support 100+ languages, and carry a 128K-token context window. Matryoshka embeddings let you pick output dimensions from 256 to 2048, and float, int8, and binary output formats cut vector storage costs. Fast delivers 2.4× higher document throughput on average, which matters more for ingestion jobs than for the live path.
Availability is via the Cohere API and Model Vault, Microsoft Foundry, and Amazon SageMaker, with self-hosting supported for both tiers.
How I’d decide
If your retrieval lives on visually rich or financial documents, the parsed-PDF and page-image numbers are the differentiators to verify against your own corpus — the shared embedding space makes it cheap to run both tiers side by side on the same index. If your workload is clean text at moderate volume, the gains over existing competitors look narrower, and the main reason to switch is operational: one index, two cost points, tunable per workload.
The catch is that RCP-nDCG@10 is new and vendor-run. Before committing an index rebuild to any tier, run your own queries against both models — the 1.6% and 2.7% cross-model losses are averages, and your corpus may not be average.
