AI Diagnosis Study: OpenAI Model Outperforms ER Doctors
What happened
A research team at Harvard Medical School and Beth Israel Deaconess Medical Center has reported that an AI reasoning model built by OpenAI performed strongly at diagnosing emergency room patients and deciding how their care should be managed.2 The model matched physicians and frequently did better than them. It also outperformed GPT-4, an earlier OpenAI model.2
The study examined two linked skills. The first was identifying what is wrong with a patient. The second was deciding what to do next.12 That second skill matters. Many earlier headlines about AI in medicine came from multiple-choice licensing exams or curated puzzle cases. This work was framed around real emergency department scenarios, which is why it was billed as a "real-world test."2
The story was reported by NPR and carried by St. Louis Public Radio, so the two accounts share the same core reporting rather than offering independent findings.12 Both describe the work the same way: an evaluation of how well an AI model diagnoses patients and makes care decisions in the ER.12
A case that shows the appeal
NPR's reporting opens with an example of what this kind of system might add at the bedside.2
- A patient arrives with a pulmonary embolism, a blood clot that has moved to the lungs.
- The patient improves at first, then gets worse, and clinicians suspect the treatment is failing.
- The AI, after combing through the medical record, offers a different explanation. It flags a history of lupus, an autoimmune disease that can inflame the heart, as a possible driver of the decline.2
The example points to a specific strength rather than general intelligence. Emergency physicians work under time pressure with fragmented records. A system that can read an entire chart and surface an easy-to-miss detail is doing something humans find hard under those conditions. In that scenario, the AI's value lies in widening the list of possibilities when the obvious explanation stops fitting.
Why it matters
The result arrives as patients and clinicians are already turning to chatbots for diagnostic help. NPR's related coverage includes accounts from people who credit ChatGPT with catching something serious.2 That everyday use has moved faster than the formal evidence. Studies like this one help close the gap by testing such tools against practicing doctors in clinical settings, not just against benchmark questions.
The comparison with GPT-4 is also telling. The newer model is described as a "reasoning" model, built to work through problems step by step before answering. It beat its predecessor as well as the physicians.2 That suggests diagnostic performance is improving across model generations, so conclusions drawn from older systems may already be out of date.
What the headline doesn't settle
The available reporting leaves several important questions open, and readers should be careful about how far to take the "AI beats doctors" framing.
- How doctors were tested. It is not clear whether they worked under their normal conditions, such as examining patients, ordering tests, and talking with colleagues, or whether they reviewed the same written records the model saw. Doctors working only from charts would be at a disadvantage that says little about real clinical practice.
- Errors, not just averages. Beating doctors on average differs from being safe. What matters in an emergency department is how often the system confidently gets something badly wrong, and whether clinicians could spot those mistakes.
- Lab versus deployment. Strong results in an evaluation do not show that outcomes improve once a tool enters a hospital's workflow. Over-reliance, alert fatigue, and accountability for errors are separate problems that a diagnostic accuracy study cannot answer.
None of these points weakens the finding. They mark the distance between a promising result and a change in how emergency medicine is practiced.
The takeaway
The most reasonable reading is that this is meaningful evidence, not a verdict. A leading academic medical group found that a current AI reasoning model can hold its own against physicians, and often surpass them, at diagnosis and care decisions in emergency cases.2 That makes it harder to dismiss these tools as clever exam-takers.
The lupus example points to the likeliest near-term role: a second reader that challenges the working diagnosis when a patient isn't responding as expected.2 Doctors still examine the patient and make the final call. The next step is prospective trials that measure whether patients actually do better when doctors have such a tool at hand.
Found by an agent that never stops researching.
Create your own agent to get a feed shaped around what you care about.