AI paper index
Safety of patient-facing AI-generated medical advice: a structured clinical evaluation using acute appendicitis scenarios
One-line summary
An AI research paper on Safety of patient-facing AI-generated medical advice: a structured clinical evaluation using acute appendicitis scenarios.
Engineering notes
Engineering notes will be added by the aipentium editorial team.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。
Original abstract
Abstract Background Patient-facing large language model–based chatbots are increasingly used to seek medical advice and support health-related decision-making. While these tools provide accessible health information, their safety in high-risk, time-sensitive clinical scenarios remains insufficiently understood. Acute appendicitis represents a clinically relevant exemplar in which inappropriate or delayed guidance may lead to serious patient harm. This study aimed to evaluate the safety of patient-facing AI-generated advice for suspected acute appendicitis and to compare major harm across question categories, AI models, prompting conditions, and repeated responses. Methods We evaluated AI-generated 540 AI-generated responses derived from a predefined factorial design comprising 30 patient-style questions, three AI systems, two prompting conditions, and three repeated responses per condition. No formal power-based sample size calculation was performed. Questions were categorized into symptom interpretation, medication use, and the safety of waiting at home. Three publicly accessible AI models —ChatGPT (GPT-5.2 Instant; GPT-5.2 released December 11, 2025), Gemini (Gemini 3 Flash; released December 17, 2025), and Microsoft Copilot (GPT-5–based automatic model routing)— were assessed with and without structured prompting, with three repeated responses generated per condition. Standard publicly deployed versions were used without investigator fine-tuning or model modification. All systems were accessed through their freely available public consumer-facing interfaces rather than APIs. Queries were conducted from December 18, 2025 to January 29, 2026. The structured prompt was investigator-developed through iterative discussion and refinement by three surgeons with experience in emergency abdominal conditions, without patient or public involvement. Responses were coded, randomized, and independently evaluated under blinded conditions by the same three surgeons, with independent expert clinical assessment serving as the reference standard for safety performance. Major harm was predefined as advice plausibly associated with delayed diagnosis, inappropriate treatment, or clinically significant patient harm. Safety outcomes were assessed using two definitions: harm identified by at least one evaluator (H_any) and harm identified by a majority of evaluators (H_maj). Major harm frequencies were compared across question categories, prompting conditions, and AI systems, and interrater agreement was assessed using Cohen’s kappa; reproducibility across repeated generations was examined using worst-case repeat analysis. Results Potentially harmful advice occurred across all question categories and was most frequent in medication-related scenarios (H_any 28.8%; H_maj 23.8%), compared with symptom interpretation (19.4%; H_maj 16.8%) and home observation/waiting questions (15.9%; H_maj 13.5%). Structured prompting did not consistently reduce major harm rates, and safety profiles were broadly similar across AI models. Worst-case repeat analyses demonstrated that potentially harmful advice could recur across repeated responses under identical conditions. Question-level analyses revealed substantial heterogeneity, with clinically nuanced scenarios—particularly those involving comorbid conditions or pregnancy—more likely to elicit majority-rated major harm. Interrater agreement was good to excellent across most evaluation domains, with lower agreement for patient comprehensibility. Conclusions AI-generated medical advice for suspected acute appendicitis carries a non-negligible risk of potentially harmful recommendations, particularly in clinically nuanced scenarios. Safety performance varies substantially according to question context and may not be adequately reflected by average performance metrics alone. Patient-facing AI tools should not be relied upon for self-management decisions in high-risk clinical situations and, if used, should be framed only as adjunctive information that does not delay professional medical evaluation.
Links and sources
Need this topic turned into a technical roadmap?
aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.
Request B2B AI research
Comments