AI for Science

Harvard ER Study: o1 Outdiagnosed Two Attending Physicians

A Harvard team publishing in Science finds o1 beat two attending physicians on ER triage diagnoses, 67% vs 50-55%. The numbers, the critiques, and what they mean for clinical AI.

Harvard ER Study: o1 Outdiagnosed Two Attending Physicians — article cover
On this page6 SECTIONS
  1. How the Study Was Run
  2. What the Numbers Say
  3. Why the Gap Peaks at Triage
  4. Critiques and Limits
  5. What It Means for Clinical and Product Teams
  6. Sources

On April 30, 2026, Science published a study from a team at Harvard Medical School and Beth Israel Deaconess Medical Center with a blunt result: on emergency-room diagnostic tasks, OpenAI’s o1 reasoning model outperformed two attending physicians. TechCrunch’s May 3 coverage pushed it into the mainstream. This is not another benchmark-flex exercise. The study ran on real electronic health records from actual ER arrivals, and the physicians scoring each diagnosis did not know whether it came from a human or a model. For teams weighing where clinical AI is actually viable, the numbers are worth reading closely — and so are the limits.

How the Study Was Run

The work was led by Arjun Manrai, who heads an AI lab at Harvard Medical School, and Adam Rodman, a physician at Beth Israel. The team took 76 patients who arrived at the Beth Israel emergency room and had o1, GPT-4o, and two internal medicine attendings produce diagnoses at the same checkpoints, drawing on the same electronic record information available at each moment. Two other attendings, blinded to the source of each diagnosis, scored the results. Nothing was pre-processed for the models. The authors’ claim is sweeping: large language models “have eclipsed most benchmarks of clinical reasoning.”

What the Numbers Say

At triage — the moment of scarcest information — o1 delivered an exact or near-exact diagnosis in 67% of cases, against 55% for one physician and 50% for the other. As clinical detail accumulated, o1 climbed to 82% while the physicians landed between 70% and 79%, a gap the study flags as not statistically significant. The treatment-planning substudy is where the gap explodes: across 5 case studies and 46 doctors, the AI scored 89% while doctors working with conventional resources, search engines included, scored 34%. In one case, o1 correctly tied a lung-clot patient’s worsening symptoms to a lupus history the human doctors had skipped past.

Why the Gap Peaks at Triage

Triage is the emergency department’s lowest-information, highest-urgency moment: the patient just arrived, labs are pending, and the clock is running. A model is relatively strong here because it can weight every line of the record simultaneously and does not drop a rare but decisive piece of history to fatigue or anchoring. Rodman told The Guardian that large language models are among “the most impactful technologies in decades” and sketched a future “triadic care model — the doctor, the patient, and an artificial intelligence system.” Manrai stayed more restrained: “I don’t think our findings mean that AI replaces doctors,” but rather “a really profound change in technology that will reshape medicine.”

Critiques and Limits

The sharpest critique came from ER physician Panthagani, who called it “an interesting AI study that has led to some very overhyped headlines.” Her core objection: the human baselines were internal medicine attendings, not emergency physicians, and an ER doctor’s job is not to guess the ultimate diagnosis but to rule out what could kill the patient — a different task by definition. Technically, the models only consumed text. They could not read the patient’s appearance, distress, or bedside cues; the authors themselves compare the AI to a second opinion derived from paperwork. The paper is explicit that the findings do not justify live ER decision-making yet and call for prospective real-world trials. Rodman added a governance gap: there is “no formal framework right now for accountability” for AI diagnoses, and patients want humans in the loop for life-or-death calls.

What It Means for Clinical and Product Teams

The adoption backdrop matters as much as the study: nearly one in five US physicians already uses AI to assist diagnosis, and in the UK, 16% of doctors use it daily with another 15% weekly. Document-dense, information-fragmented clinical reasoning is now a place where consumer-grade models slot in naturally. Three takeaways for builders. First, triage and differential diagnosis are the highest-value early touchpoints because the input is almost entirely text. Second, the blinded scoring and the “exact or near-exact” grading rubric are pragmatic designs you can lift directly into your own evaluation pipeline. Third, accountability and specialty matching are the questions you must answer before deployment — the internal-medicine-baseline controversy is a textbook case study in study design for the next iteration. The paper’s central claim is not that AI is ready to run the ER; it is that AI is good enough to warrant clinical testing.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL