The humble stethoscope has been a physician’s companion for more than two centuries, but the skill of interpreting what it hears, the soft whoosh of blood through heart valves, the crackle of fluid in diseased lungs, has always depended on years of human training. Now, a pair of researchers at Saudi Electronic University in Riyadh

Google AMIE study adds evidence of medical AI’s promise with patients, under doctors’ supervision

[Adobe Stock]
Since early 2023, large language models (LLMs) have posted passing-level scores on medical licensing exam questions. A study published that February found that the first iteration of ChatGPT, based on the GPT-3.5 model that debuted in 2022, scored at or near the passing threshold on publicly available questions from all three steps of the United States Medical Licensing Examination (USMLE), and a few months later Google’s Med-PaLM 2 reached 86.5% on USMLE-style multiple-choice questions. Performance has been far less predictable once people use the models. In a randomized Oxford experiment in which members of the public worked through written medical scenarios, LLMs tested alone identified the relevant condition in 94.9% of the scenarios, but participants using those same models did so in fewer than 34.5% of cases, no better than people who used any source they chose. A Mount Sinai study found models changed clinical recommendations when only a patient’s demographic details changed.
The dynamic is steadily shifting with the evolving capabilities of LLMs. The most recent data point comes via a study published Oct. 8 in The Lancet. It moves the test out of the exam room. Google and Beth Israel Deaconess Medical Center had primary-care patients chat with AMIE, Google’s conversational diagnostic system, before urgent visits while a physician watched every exchange. The authors report that none of the conversations had to be stopped. Alphabet funded the study, and study author Adam Rodman of BIDMC was a visiting researcher at Google during part of it.
In the study, AMIE took each patient’s history before the visit. In the 44 cases where physicians reviewed the AMIE transcript beforehand, they said it helped them prepare 75% of the time and may have changed how they handled the appointment 57% of the time, according to BIDMC. Among the 98 patients who completed both the chat and the appointment, AMIE’s first seven diagnostic suggestions included the final diagnosis in 90% of cases; its top suggestion matched in 56%. Supervising physicians identified one hallucination and, separately, added clinical clarification in five cases.
Patients’ attitudes toward AI in health care also improved after the chats and stayed elevated after they saw their doctor in person.
The study is small, involving a single clinic, 98 patients, and a physician in the loop. But its evidence also points to a continued shift in LLMs’ capabilities in medicine.
That shift has come in overlapping stages.
Phase 1: Big promises, earned skepticism
Through the 2010s, medical AI largely meant narrow tools, and the marketing ran well ahead of them. Deep-learning systems learned to read specific kinds of images, and in April 2018 the FDA authorized IDx-DR, which screens retinal photos for diabetic retinopathy, as the first autonomous AI diagnostic device. Broader ambitions fared worse. IBM pitched Watson, fresh off its 2011 Jeopardy! win, as a cancer-treatment adviser. Despite considerable investment, one of its core healthcare partners, MD Anderson, would eventually shelve its Watson project after spending more than $62 million. Internal IBM documents reported by STAT in 2018 showed the system recommending “unsafe and incorrect” treatments. IBM ultimately sold healthcare data and analytics assets from its Watson Health business in 2022.
Throughout the Watson saga, research into deep learning for medical images kept advancing. A 2016 JAMA study from Google showed an algorithm could detect diabetic retinopathy in retinal photos, and a 2017 Nature paper from Stanford reported a network that classified skin cancers about as well as board-certified dermatologists.
While ChatGPT helped put chatbots on the map, research in the late 2010s explored the potential of earlier chatbot systems. In 2018, Babylon Health claimed its chatbot outscored the average candidate on a UK GP licensing exam; a Lancet review found no convincing evidence that it performed better than doctors, and the company collapsed in 2023. Even AI’s champions overreached: Geoffrey Hinton said in 2016 that people should stop training radiologists, a prediction he has since called too hasty. If anything, the numbers are pointing in the opposite direction. Mayo Clinic alone grew its radiology staff from about 260 in 2016 to more than 400 while adopting hundreds of AI models, and researchers at the American College of Radiology’s Neiman Health Policy Institute project that the U.S. radiologist shortage could last into the 2050s.
Given the hype of this phase, a counterweight of skepticism emerged. Forecasts such as venture capitalist Vinod Khosla’s 2012 prediction that technology would replace 80% of what doctors do met a cooler reception in the clinic. In a 2018 survey of 720 UK GPs, 68% considered it unlikely that technology would ever fully replace physicians in diagnosis.
Phase 2: Useful, but not yet reliable
ChatGPT’s arrival in late 2022 traded narrow focus for breadth. Instead of one model trained per task, a single general-purpose system could field questions across specialties, from exam vignettes to patient messages, albeit with uneven reliability. The licensing-exam scores offered early evidence of LLMs’ emergent breadth of medical knowledge. Within months, a JAMA Internal Medicine study found that evaluators preferred chatbot answers to patients’ online questions over physicians’ answers in 78.6% of cases, and rated them more empathetic.
A jack-of-all-trades-master-of-some dynamic was also present in the early days of LLMs. For instance, a 2023 New England Journal of Medicine report by Microsoft researchers on GPT-4 described useful medical applications alongside fabricated information, delivered with the same confidence as correct answers. Physicians who tested the models early saw both sides. Isaac Kohane, a physician who chairs Harvard Medical School’s biomedical informatics department, wrote in a 2023 book that GPT-4 performed clinically “better than many doctors I’ve observed.” Cardiologist Eric Topol, reviewing the book, also cautioned that “there is a big problem with hallucinations, which hasn’t materially changed with GPT-4,” and that the models “have not been validated in the real world of health systems.”
Phase 3: From exam questions to patient conversations
AMIE’s path to the Lancet trial ran through a simulation. In a 2025 Nature study, Google pitted an earlier version of AMIE against 20 primary-care physicians across 159 case scenarios played by patient actors. AMIE beat the physicians on diagnostic accuracy and on most consultation-quality ratings. But everyone typed through a text-chat interface and the patients were trained actors, so its performance in clinical practice went untested. The Lancet study was built to close that gap, putting AMIE in front of patients before scheduled appointments.
Strong model performance raised the practical question of how clinicians could use it effectively. In Goh and colleagues’ 2024 randomized study, GPT-4 alone outperformed participating physicians on a diagnostic-reasoning rubric. Giving physicians access to it produced no statistically significant improvement over conventional resources.
Phase 4: An AI second opinion
The AMIE trial may also offer a glimpse of where patient-facing AI is headed. At a January event tied to the J.P. Morgan Healthcare Conference, Anthropic CEO Dario Amodei described how his sister, co-founder Daniela Amodei, developed an infection while pregnant. “She went to a bunch of fancy doctors, and for some reason, they all thought that this was a viral infection,” he said, as we reported earlier this year. After she uploaded her medical data, “Claude basically gave a second opinion where Claude said, ‘I think this is a bacterial infection.’” Amodei framed it in those terms: “It’s a second opinion, and that is usually very helpful,” he said.
Research increasingly backs that narrower claim. Probabilistic as they are, LLMs can inform clinical decisions. In a 2025 Nature study of 302 challenging published cases, AMIE included the correct diagnosis among its top 10 suggestions 59.1% of the time, compared with 33.6% for 20 unassisted clinicians. The models remain uneven, though. As the Oxford and Mount Sinai studies showed, they can falter when laypeople drive the conversation or when only a patient’s demographic details change.
