Large Study Finds AI Patient‑Portal Replies Often Misaligned With Doctors
Dartmouth researchers show personalization reduces clinician editing without removing the need for human review.
Overview
- A Dartmouth paper published at the Association for Computational Linguistics in July 2026 analyzed 146,000 real patient‑physician portal messages from 10,105 patients to test how well AI drafts match clinician replies.
- Researchers tested six commercial large language models, including ChatGPT, Claude and Gemini, and found drafts frequently omitted key follow‑up questions, added irrelevant or inaccurate medical details, and produced overly long replies.
- The flaws can increase clinician work because doctors may spend more time correcting AI drafts than writing messages from scratch, and prior independent research found a measurable share of AI replies could pose serious safety risks.
- The team developed a personalization method called TADPOLE and reported roughly 33% better alignment with physicians’ style and about 26% less editing when models were adapted to individual clinicians.
- The study has heightened scrutiny as health systems expand AI use; authors urge continued clinician oversight, further testing of real editing time and user acceptability, and careful policy choices during rollouts such as the NHS deployment.