Unbiased headlines. Facts, not spin.
Every story is an unbiased news briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
Harvard Study Found OpenAI's o1 Model Diagnosed Cases Better Than Physicians in Head-to-Head Tests

Researchers at Beth Israel Deaconess Medical Center and Harvard Medical School put an OpenAI model up against practicing physicians on real diagnostic cases. The AI won, and not by a little.
In one test using case studies, the o1 model generated a list of possible diagnoses and landed on the right one 78% of the time. Physicians working the same cases got it right about 30% of the time, according to the study published April 30 in the journal Science.
"We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines," said Harvard Medical School professor Arjun K. Manrai, one of the study's senior authors.
A second test gave o1 five clinical vignettes pulled from real patient cases and asked what to do next. Two physicians graded the answers blind. The model averaged 89%. Human doctors given the identical test averaged 34%.
The researchers also ran an emergency room simulation, where information is often incomplete and decisions have to happen fast. The model was scored against two attending physicians, with two other attending physicians grading both sets of answers without knowing which came from the machine.
"Overall, o1 outperformed both [an earlier LLM] and two expert attending physicians, as assessed by two other attending physicians who both were blinded to the source of the differential diagnosis," the study states. At the initial triage stage specifically, o1 identified the exact or a very close diagnosis in 67.1% of cases. The two human doctors managed 55.3% and 50%, respectively.
The o1 model tested here is an OpenAI preview tool that has since been superseded by o3, meaning the model that beat doctors in this study is already outdated technology by OpenAI's own release cycle. Whatever gap existed in this study, current models are a generation past it.
The Harvard results follow other recent studies showing similar gains. A Swedish study published in The Lancet in January found AI-assisted mammogram reading outperformed standard screening at catching breast cancer. Separately, researchers at the Mayo Clinic led by Sovanlal Mukherjee built a model that scanned abdominal CT images from patients later diagnosed with pancreatic cancer. The AI flagged the disease an average of 475 days earlier than clinicians did, and in some cases up to three years earlier. That study ran in the journal Gut.
"Attaining such early detection would substantially augment the probability of cure and improved survival," the Mayo Clinic researchers wrote. Pancreatic cancer is one of the deadliest cancers precisely because it's usually caught late. A model that spots it more than a year sooner is not a marginal improvement, it's the difference between treatable and not.
None of this means software is replacing your doctor next year. Peter Brodeur, one of the Harvard study's authors, put it plainly in a press release: "humans should be the ultimate baseline." The study measured diagnostic reasoning on written cases, not the messier reality of an actual patient encounter, physical exams, bedside judgment, or the liability and trust questions that come with letting an algorithm make the call.
There's also a fair concern buried in all this enthusiasm: these are benchmark tests, not real-world deployment data. A model acing case studies graded by physicians is not the same as a model performing safely across millions of unpredictable patients with messy histories, drug interactions, and incomplete records. Skeptics of rushing AI into clinical practice have a legitimate point that the gap between a controlled study and a hospital floor is real, and that overconfidence in AI diagnostics before that gap is closed could hurt patients.
Some states are moving in the opposite direction from where the data points. Nevada now bans AI systems from saying anything that "implicitly indicates" they're "capable of providing professional mental or behavioral health care," and bars them from providing any service that would "constitute the practice of professional mental or behavioral health care." Illinois has passed a similar restriction.
Those laws are aimed at protecting patients from bad actors selling unproven chatbot therapy, a legitimate worry given how many mental health apps have made unverified claims. But they're written broadly enough to also block the kind of diagnostic assistance the Harvard, Mayo Clinic, and Lancet studies suggest could actually save lives, especially in specialties like radiology and oncology where speed of detection is everything.
The unresolved question is whether regulators can write rules that stop the genuine bad actors in AI mental health services without also freezing out diagnostic tools with published, peer-reviewed results behind them. Right now, Nevada and Illinois have chosen a blanket restriction over that distinction. Whether other states follow that model, or instead build frameworks that let hospitals use tools like o1 or the Mayo Clinic's pancreatic cancer model under physician supervision, will likely be decided state by state and won't be argued out on scientific merit alone.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.