If you build products where an agent has to find everything meeting a condition — every qualifying vendor, every affected location, every person with a given credential — the uncomfortable question is how much your search backend is actually costing you in missed results. Exa published a preview of ATLAS on October 8, 2026, a benchmark built to answer that question, and the early numbers are sobering: no agent run under $1 per task scored above 0.5 row F1, and even the most expensive systems missed roughly a third of the golden results.
Why existing search evals go stale
Exa’s argument is that the popular agentic-search benchmarks have a shelf-life problem. Per their table, BrowseComp, WideSearch, and DeepSearchQA are 48–61% memorized — meaning at least one frontier model can recall the full answer set closed-book — and sit close to saturation. Newer LLM-judged benchmarks like Perplexity’s WANDR stay fresh but cost $18k–50k to grade ten runs, versus about $2 for ATLAS, and are sensitive to grader design choices.
The practical consequence for builders: a benchmark that agents can partially answer from memory isn’t measuring your search stack. Exa demonstrated this directly — when they degraded search by hiding the top 7 of every 10 results, their searcher kept 81–89% of its score on the older benchmarks but lost roughly half its score on ATLAS.
How the tasks are built
ATLAS contains 547 tasks, each generated from clusters of anonymized real search demand. A task asks the agent to discover every entity meeting precise conditions, then enrich each with 2–10 multi-hop attributes. Construction is automated: two agents from different model families (Codex with GPT-6 Astra, Claude Code with Opus 5.5) build each golden table, spending a median of 8 agent-hours and about 1,200 searches, with independent verification and targeted repair stages. Every agent searches through a provider-neutral tool that interleaves Exa, Brave, and Perplexity results with provider names hidden, so the gold isn’t biased toward one index.
Because the pipeline is automated, the dataset can be regenerated from recent search trends — which is how Exa plans to fight memorization over time. Disputed cells are left ungraded (7.1% of cells), and 3.8% are verified blanks, testing whether agents can correctly conclude that information doesn’t exist.
What the early runs show
Three findings stand out from Exa’s runs:
- Cheap and complete don’t coexist yet. Max-effort agents beat low-compute ones consistently, and nothing under $1/task crossed 0.5 row F1.
- The search backend itself moves the needle. With the model harness fixed, switching search backends shifted scores by 16%, with Exa claiming the cost-performance Pareto frontier.
- Width and depth are real difficulty multipliers. 74% of tasks need 10+ entities, and full discovery plus enrichment spans a median of 18 domains per task. On WideSearch, a strong agent solved 48% of tasks citing a single website; on ATLAS, that figure is 0.2%.
If your product leans on entity discovery and enrichment — the same pattern behind structured data agents like the ones I wrote about in place data your agent can act on — ATLAS is worth studying as a spec for what “good” looks like. If your agent just needs one right answer, the older benchmarks may still tell you what you need.
The takeaway for builders
The stated challenge ATLAS poses is efficient search: solve each task in under a minute and under $1. No current system gets close. Treat that as your planning assumption — if completeness matters to your product, budget real compute per query, and when evaluating search providers, run your own ablation with the harness held fixed rather than trusting any single vendor’s leaderboard. Note this preview is based on Exa’s own runs of its own benchmark, so the Pareto-frontier claim deserves independent replication before you make purchasing decisions on it.
