While a high-profile study suggests algorithmic models significantly outperform human practitioners in clinical diagnostics, experts warn the findings are largely observational and fail to account for the algorithm's lack of a medical degree.
The peer-reviewed paper, discussed in this week's STAT Health Tech newsletter and published in the New England Journal of Medicine, found that a large language model correctly identified complex pathologies in 99.8% of test cases, compared to a 73% success rate among board-certified internists. However, lead authors were quick to note that the study’s methodology—which relied on feeding the AI millions of verified patient charts—presents severe limitations. Because the data was evaluated retrospectively, experts caution it is entirely possible the AI’s flawless performance is merely a correlative anomaly rather than evidence of a causal relationship between the software and accurate medical care.
Furthermore, industry observers have flagged potential confounding variables in the study design. Authors of the paper disclosed that the AI was not subjected to the sleep deprivation, administrative burnout, or pharmaceutical rep lunches that typically inform real-world clinical decision-making. A secondary analysis of the control group revealed that human doctors were significantly better at prescribing whatever antibiotic they had most recently seen on a branded pen, a critical metric the study's authors inexplicably omitted from the primary endpoints.
While the initial data appears robust, we must be incredibly careful not to conflate the algorithm's unprecedented ability to instantly identify rare pathologies with actual medical expertise.
Dr. Thorne added that until researchers can conduct a 40-year, double-blind, placebo-controlled trial confirming the AI is not simply making incredibly lucky, highly-educated guesses every single time, patients should continue waiting six weeks for an in-person consultation.
The STAT report also highlighted ongoing debates over telehealth's role in abortion access, another area where researchers warn AI integration remains "statistically immature" due to the software's inability to project moral judgment through a video portal. At press time, the American Medical Association advised the public to treat the AI's life-saving, entirely accurate diagnoses as strictly off-label.