An AI Model Matched or Beat Doctors on Real ER Cases
In a study published in Science, OpenAI's o1-preview matched or beat internal medicine physicians on 76 real, unedited emergency room cases, with the biggest gap in its favor at initial triage, when the least information is available. The model wasn't tested against emergency medicine specialists, and this was a retrospective study, not a live trial.
Researchers at Harvard and Beth Israel Deaconess wanted to know if an AI reasoning model could handle a doctor's messiest job: reading real, unedited patient records and figuring out what's wrong. In a study published April 30 in Science, the model matched or beat physicians at every stage of real emergency room cases.
The 30-second summary
- What happened? Researchers tested OpenAI's o1-preview against two internal medicine attending physicians on 76 real, unedited emergency room cases from a Boston hospital, at three points: initial triage, early assessment, and hospital admission. Two other doctors, not told which diagnosis came from a human or the AI, judged the results blind. The AI matched or beat the physicians at every stage.
- Why does it matter? This is one of the largest tests of an AI model on real, messy hospital data rather than cleaned-up textbook cases, and the gap favored AI most when the least information was available, exactly the moment early triage decisions get made.
- What is the catch? The comparison was against internal medicine physicians, not emergency medicine specialists, and this was a retrospective study, not a live trial where the AI actually guided patient care.
KEY NUMBER
In a separate part of the study, using published diagnostic case reports, o1-preview reached the correct or very close diagnosis in 88.6% of cases, compared with 72.9% for the earlier GPT-4 model.
What Happened
The team, led by physicians and computer scientists at Harvard Medical School and Beth Israel Deaconess, ran six separate experiments comparing OpenAI's o1-preview, a reasoning model that works through a problem step by step before answering, against physicians and older AI models. The core test used 76 real emergency department cases, presented exactly as they appeared in the electronic health record, unedited and full of the noise a real chart contains.
At three points, initial triage, early assessment, and hospital admission, the model's diagnosis was compared against two internal medicine attending physicians working from the same notes. Two separate attending physicians then judged the results without knowing which diagnosis came from a person and which from the AI. The model matched or beat the physicians at every stage, with the biggest gap at triage, the point where the least information is available and clinicians can least afford to be wrong.
Why It Matters
Most earlier tests of medical AI relied on cleaned-up textbook cases or multiple-choice exams. This study deliberately used unedited hospital records instead, and the model still performed well. But the comparison has a real limit worth naming: the physicians in the study were internal medicine specialists, not emergency medicine doctors, the kind who actually staff ER triage desks. Emergency physician Kristen Panthagani, who was not involved in the study, pointed out that this mismatch means headlines calling it a win against "ER doctors" overstate what was actually tested.
Before We Overstate the Result
- This was a retrospective study comparing outputs on existing records, not a live trial where the AI guided real patient care.
- The physicians used as the comparison group were internal medicine attendings, not emergency medicine specialists.
- On identifying rare but critical "cannot-miss" diagnoses, the model was not significantly better than physicians, the exact category where a missed AI error could be most dangerous.
- Every group in the study, including the AI, overestimated how likely rare conditions were once test results came back, a known bias none of them escaped.
What Happens Next
The study's authors, including co-senior author Adam Rodman, are explicit that this doesn't mean AI is ready to diagnose patients on its own. Rodman has said he's uneasy about how the results could be used, and pointed out there's no formal framework yet for who is accountable when an AI diagnosis is wrong. The researchers are calling for prospective clinical trials, ones where the AI's suggestions actually influence real patient care under supervision, before any of this moves from research finding to hospital practice. It joins a broader wave of hospital-focused AI research, including a separate model built to warn clinicians about kidney injury before the damage becomes obvious.
Takeaway
The clearest signal in this study isn't that AI can win a diagnostic contest, prior research already suggested that. It's that the model held up on real, unedited hospital data instead of cleaned-up textbook cases, and did best exactly when doctors have the least information to work with. Whether that translates into safer care, rather than just a better benchmark score, is the prospective trial nobody has run yet.
Verified topics and entities
Sources and citations4 sources
External references used to support the reporting in this article.
- Performance of a large language model on the reasoning tasks of a physician
- Landmark test of clinical reasoning finds AI outperformed physicians, raising bar for more serious testing
- An AI model beat doctors at diagnosing patients, in a new study
- In Harvard study, AI offered more accurate emergency room diagnoses than two human doctors
Published by
NewTqnia Artificial Intelligence Desk
An institutional editorial team within NewTqnia