By Brian Buntz
Publication Date: 2026-10-09 23:20:00
[Adobe Stock]
Since early 2023, large language models (LLMs) have posted passing-level scores on medical licensing exam questions. A study published that February found that the first iteration of ChatGPT, based on the GPT-3.5 model that debuted in 2022, scored at or near the passing threshold on publicly available questions from all three steps of the United States Medical Licensing Examination (USMLE), and a few months later Google’s Med-PaLM 2 reached 86.5% on USMLE-style multiple-choice questions. Performance has been far less predictable once people use the models. In a randomized Oxford experiment in which members of the public worked through written medical scenarios, LLMs tested alone identified the relevant condition in 94.9% of the scenarios, but participants using those same models did so in fewer than 34.5% of cases, no better than people who used any source they chose. A Mount Sinai study found models changed clinical recommendations when only a patient’s…

