Would You Trust a Robot Over Your Doctor?
A Harvard–Stanford study found that an AI reasoning model outperformed emergency room doctors in diagnostic accuracy using real patient cases. But while the results are impressive, the bigger question remains: who is accountable when AI gets a diagnosis wrong?

AI vs. Emergency Room Doctors: What the Research Found
Would You Trust a Robot Over Your Doctor? A Harvard and Stanford research team just tested an AI reasoning model against real emergency room doctors, using actual patient cases from a Boston hospital, unedited and messy the way real medical records actually are. At the triage stage, the point where a doctor has the least information and the most pressure to get it right, the AI landed on the correct or near correct diagnosis 67.1 percent of the time. The two attending physicians it was compared against scored 55.3 percent and 50 percent (Brodeur et al.). That gap did not shrink much as more information came in either. By the time a patient would normally be admitted, the model was still ahead.
Why These Results Are Surprising
I will be honest, my first reaction to that stat was denial. Doctors train for over a decade. They do residencies, they see thousands of patients, they build the kind of instinct that is supposed to come from experience, not from a chatbot scanning text. But the study did not cherry pick an easy comparison. The researchers specifically avoided cleaning up the data because they wanted to see how the model performed with the same messy, incomplete information a real doctor gets handed at 3 a.m. in an ER. It still won.
AI Is Not Replacing Doctors
Here is the part people keep skipping past when they share this story. None of the researchers are saying AI should replace physicians. One of the lead authors, Adam Rodman, is a practicing doctor himself, and he has said publicly that there is currently no formal system for holding an AI accountable if it gets a diagnosis wrong, and that patients still want an actual human guiding them through a life or death decision. That matters more than the accuracy number. A tool being right more often than a tired physician on a night shift is not the same thing as that tool being ready to carry legal and emotional responsibility for someone's life.
The Real Question: Who Is Accountable?
So where does that leave us. I think the honest answer is that the technology has quietly gotten ahead of the systems built to regulate it. If an AI model is already outperforming trained doctors on diagnostic accuracy, the conversation cannot just be about whether the technology works. It has to be about who is accountable when it is used, and right now nobody has actually answered that question.
Works Cited
Brodeur, Peter G., et al. "Performance of a Large Language Model on the Reasoning Tasks of a Physician." Science, vol. 392, no. 6797, 2026.
Frequently Asked Questions
Share this post
Comments (0)
No comments yet. Be the first to comment!
