A new Harvard-led study is adding more evidence that advanced AI models can perform well on some medical reasoning tasks, including diagnosis exercises built from real emergency department cases.
The study, published in Science, came from researchers at Harvard Medical School and Beth Israel Deaconess Medical Center. The team tested OpenAI models across several clinical settings, including a comparison involving emergency department records from Beth Israel.
The headline result is attention-grabbing: OpenAI’s o1 model produced the exact or very close diagnosis more often than two internal medicine attending physicians at the earliest point in the emergency care process. But the details are important. The study was not a trial of AI running an emergency room, and it did not show that patients should rely on chatbots instead of doctors.
It showed something narrower, and still meaningful: when given the same text-based medical record information available at specific points in care, one reasoning model performed at or above the physician comparison group on diagnostic accuracy.
What the Harvard Team Tested
In one part of the research, investigators looked at 76 patients who had come through the Beth Israel emergency department. The study compared diagnoses written by two internal medicine attending physicians with diagnoses generated by OpenAI’s o1 and GPT-4o models.
The diagnoses were then reviewed by two other attending physicians who did not know whether each answer came from a human doctor or an AI model. That blinded review was meant to reduce bias in judging the quality of the responses.
The researchers evaluated the model and physician diagnoses at several points in the patient’s emergency care timeline, including the initial triage stage. That early moment is especially difficult because clinicians have limited information and must make decisions quickly.
According to the study, o1 was either roughly comparable to or better than the two physicians and GPT-4o across those diagnostic touchpoints. The largest gap appeared at initial triage, when the patient information was thinnest.
Harvard Medical School’s summary of the work emphasized that the researchers did not clean up or reorganize the patient data before giving it to the models. The AI systems received the same information available in the electronic medical record at the time of each diagnostic step.
The AI Revolution in Medicine: GPT-4 and Beyond
Readers who want more background on how large language models may fit into clinical work may find this book useful. It focuses on medical uses of GPT-4 while also addressing risks and guardrails.
At triage, o1 returned the exact or a very close diagnosis in 67% of cases. The two physicians reached that mark 55% and 50% of the time, respectively.
Why the Result Is Significant
The study matters because it moves beyond carefully written medical exam questions and into messier hospital records. Real clinical data is often incomplete, uneven, and filled with distractions. A model that can reason through that kind of information could eventually become useful as a second reader for clinicians.
That does not mean the model is ready to make treatment decisions. It does suggest that AI may be able to help surface plausible diagnoses earlier in a care process, especially when doctors are working with limited time and scattered information.
The researchers framed the result as a reason to move toward prospective trials, meaning studies that test AI tools in real patient care settings under controlled oversight. That is a very different claim from saying an AI model should be deployed on its own.
The distinction matters because diagnosis is only one part of emergency medicine. Doctors also examine patients, interpret behavior and distress, assess visual cues, weigh risk, communicate uncertainty, and decide what must be ruled out immediately. A text-only model is not doing all of that.
The Study’s Limits Matter
Several caveats keep the findings from supporting the more dramatic version of the story.
First, the physician comparison group in the emergency department diagnosis experiment consisted of internal medicine attending physicians, not emergency medicine physicians. That difference is not a technicality. Emergency doctors are trained to prioritize immediate threats, stabilize patients, and decide what cannot be missed, even when the final diagnosis remains uncertain.
Emergency physician Kristen Panthagani argued that this point has been lost in some coverage of the study. Her criticism is that comparing an AI tool to physicians outside the relevant specialty does not answer the question many headlines imply: whether AI can outperform emergency doctors at emergency medicine.
Deep Medicine by Eric Topol
This book is a useful companion for readers interested in how AI could change diagnosis and medical practice without removing the human role in care.
Second, the task focused on identifying the final or near-final diagnosis. In an emergency department, the first goal is often not to name the final condition perfectly. It is to determine whether the patient might have something life-threatening and what needs to happen next.
That distinction changes how the result should be understood. A model that guesses the eventual diagnosis well may still be incomplete as a tool for front-line emergency care if it does not reliably prioritize dangerous possibilities, understand uncertainty, or fit into a safe clinical workflow.
Third, the models were tested on text-based information. The researchers noted that current foundation models can be weaker when reasoning over nontext inputs. In real emergency care, doctors use imaging, vital signs, physical exam findings, patient appearance, tone, timing, and many other signals that may not be captured cleanly in a record.
What Comes Next for AI Diagnosis Tools
The study adds to the case that reasoning models could become part of clinical decision support, especially in settings where doctors need fast second opinions or help sorting through complex records.
But the path from benchmark performance to bedside use is long. Hospitals would need clear rules for accountability, safety testing, workflow design, patient consent, privacy, and clinician oversight. Adam Rodman, one of the study’s lead authors and a physician at Beth Israel, has also cautioned that medicine still lacks a formal accountability framework for AI-assisted diagnoses.
That gap is not abstract. If an AI system suggests the wrong diagnosis, misses a dangerous condition, or influences a doctor’s decision in a harmful way, hospitals and regulators will need clear answers about responsibility.
For patients, the practical takeaway is more restrained than the headline numbers. This study does not show that people should use AI instead of seeking emergency care. It shows that one advanced model performed impressively on a narrow diagnostic task using medical record text from real cases.
The most likely near-term role for systems like o1 is not replacing physicians. It is helping clinicians think through difficult cases, check their assumptions, and spot possibilities that might otherwise be missed. Even that role will need careful real-world testing before it can be trusted in high-stakes care.
The result is impressive because it is specific. It is also easy to overstate. The Harvard study suggests AI diagnosis tools are becoming more capable, but it also underscores why medicine cannot treat a strong benchmark result as a finished safety case.
