This transcript has been edited for clarity. Welcome to Impact Factor, your weekly dose of commentary on a new medical study. I’m Dr F. Perry Wilson from the Yale School of Medicine. There’s a new genre of medical papers in the “AI in medicine” space, and, like Mulder from The X-Files, I want to believe. The theme of these papers is something like “LLMs aren’t going to replace good doctors,” and that is very reassuring for this doctor, who definitely does not want to be replaced by a friendly AI. But if I’m honest with myself, my belief in human supremacy here is being shaken. This week, a new paper appeared in JAMA Network Open evaluating the performance of various large language models on a diagnostic task. It was chock full of phrases like “our evaluation suggests that despite rapid advances in pattern recognition and knowledge retrieval, current LLMs still lack the reasoning processes needed for safe clinical use.” That’s reassuring. It also states that “the promise of LLMs in clinical medicine lies in their potential to augment — not replace — physician reasoning.” Perfect. I’m happy to be augmented. I would rather not be replaced. But then I read through the study and, frankly, I’m more worried now than ever. Let me break it down for you. Researchers evaluated 21 off-the-shelf large language models, including all the major players (ChatGPT, Claude, DeepSeek, Grok, Gemini), across 29 clinical vignettes from the MSD manual. The key innovation of this study over previous evaluations is how these vignettes are structured. Other studies evaluating LLMs often present an entire case and simply ask “what’s the diagnosis?”, but that’s not how these clinical vignettes work. Instead, they develop iteratively. You get an initial presentation and then formulate a differential diagnosis, choose the appropriate tests,