Benchmarks

Epoch AI: Chinese Models Trail US Frontier by Seven Months

Epoch AI's ECI index puts Chinese models seven months behind the US frontier on average since 2023, ranging four to fourteen months. How it is measured, and the open-weight catch.

Epoch AI: Chinese Models Trail US Frontier by Seven Months — article cover
On this page6 SECTIONS
  1. How the ECI Computes “Seven Months”
  2. Four to Fourteen Months: The Shape of the Gap
  3. The Open-Weight Confound
  4. Do Seven Months Actually Matter
  5. What It Means for Developers
  6. Sources

On January 2, 2026, Epoch AI published a data analysis by Luke Emberson with a one-line conclusion: measured by the institute’s Epoch Capabilities Index (ECI), every frontier AI model since 2023 has been built in the United States, while Chinese models have trailed the US frontier by seven months on average, with the gap ranging from four to fourteen months. TIME’s graphical rundown of the US-China AI race cites the figure directly, and it is on its way to becoming a standard reference in policy arguments.

The number matters less for how it sounds than for what it is: the first public, day-by-day curve that turns “how far behind is China, exactly” into a tracked quantity. For model selection and for policy debates alike, both the measurement and its limits are worth unpacking.

How the ECI Computes “Seven Months”

The ECI is Epoch’s benchmark-based composite of model capability. The procedure: for each day, take the best-scoring Chinese model’s ECI, then look back to find how long ago a US model held an equal or lower score; that interval is the lag, with scores within one ECI point treated as equivalent. Each country’s first model — LLaMA-65B for the US, Baichuan1-7B for China — was excluded as likely not genuinely at the frontier. The dataset ships under a Creative Commons BY license, with the underlying CSV refreshed on January 6; this is built to be tracked over time.

Four to Fourteen Months: The Shape of the Gap

Two anchors stand out in the results. In May 2024, the first Chinese model to surpass GPT-4’s ECI arrived — fourteen months late, the widest gap in the observation window. And as of publication, no Chinese model had surpassed OpenAI’s o3, released in April 2025. The leading Chinese entries in the dataset include DeepSeek-R1, DeepSeek-V3.2, Qwen3-235B, Qwen3-Max, Kimi K2 Thinking, and Zhipu’s GLM series; the US frontier has rotated through GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5.

The Open-Weight Confound

Epoch flags its most important limitation itself: the country gap almost perfectly mirrors the proprietary-versus-open-weight gap. Nearly every top Chinese model ships with open weights, while US frontier models are all closed. What the curve may be capturing, then, is a business-strategy difference — open versus closed — rather than a pure capability difference. It echoes the open-source route split we covered in our 2026 opening outlook: the companies leading on raw capability and the companies leading on open weights are no longer the same set, and the two coordinate systems give different answers to “who is ahead.”

Do Seven Months Actually Matter

The Hacker News thread — 87 comments and counting — rehearses every objection. One camp attacks the metric: benchmarks leave room for “benchmaxxing,” and a score lead is not a real-world capability lead. Some argue Chinese models are distillations of US frontier models, to which others reply that DeepSeek’s MLA architecture and MoE load-balancing are widely recognized original contributions. The other camp questions whether the conclusion matters at all: seven-month-old state of the art is good enough for most real tasks, and in the open-weight and small-model market the Chinese camp is the dominant force — Alibaba’s closed flagship Qwen 3 Max is nearly invisible in Western discussion. The side effects of export controls also come up: chip restrictions have pushed Chinese teams to squeeze efficiency as far as it will go.

What It Means for Developers

Three practical conclusions. First, treat “seven months” as a baseline, not a law: the closed frontier leads the best open weights by roughly half a year, and whether your product genuinely needs that half-year determines your cost structure. Second, before reading any country-level capability comparison, read the metric definition — Epoch’s method is fully public, but swap in a different index and the conclusion can flip. Third, the number has already entered mainstream policy and media discourse (TIME’s US-China race graphs are one example) and will only appear more often; understanding how it was measured beats memorizing the digit.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL