AI paper index

Comparison of ChatGPT, Gemini, and DeepSeek responses to questions on medication-related osteonecrosis of the jaw (MRONJ): a comparative analysis of expert-rated clinical quality, readability, response time, and clinical safety risk

2026-07-25 · BMC Oral Health

One-line summary

An AI research paper on Comparison of ChatGPT, Gemini, and DeepSeek responses to questions on medication-related osteonecrosis of the jaw (MRONJ): a comparative analysis of expert-rated clinical quality, readability, response time, and clinical safety risk.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

Medication-related osteonecrosis of the jaw (MRONJ) is a clinically significant complication associated with antiresorptive and antiangiogenic therapies, and its prevention and management depend on reliable, guideline-concordant information. Large language models (LLMs) are increasingly used to generate medical information, but their performance in MRONJ remains insufficiently investigated. This study compared the responses generated by ChatGPT- 5.4 Thinking, Gemini 3 Deep Think, and DeepSeek-V3.2 in thinking mode to MRONJ-related open-ended questions in terms of clinical quality, readability, and response time. In this cross-sectional comparative study, a standardized 30-item question set was developed from the MASCC/ISOO/ASCO Clinical Practice Guideline and the Italian Consensus Update. Each question was entered separately into each model under standardized conditions, yielding 90 responses. Outputs were anonymized and independently evaluated by three blinded oral and maxillofacial surgeons using a modified Global Quality Scale on a five-point Likert scale. A post hoc Clinical Risk Classification was additionally performed to distinguish minor informational deficiencies from omissions or inaccuracies with greater potential clinical consequences. Readability was assessed using five established English-language indices, and response time was measured with an online digital stopwatch. Appropriate statistical tests were applied, and inter-evaluator reliability was assessed using the intraclass correlation coefficient. Inter-evaluator agreement was high across all models (all p < 0.001). Mean expert quality scores differed significantly among models ( p < 0.001). Post hoc comparisons showed that ChatGPT-5.4 Thinking received significantly lower expert ratings than both DeepSeek-V3.2 and Gemini 3 Deep Think, whereas no statistically significant difference was detected between DeepSeek-V3.2 and Gemini 3 Deep Think. Response time also differed significantly ( p < 0.001): DeepSeek-V3.2 produced the fastest responses, Gemini 3 Deep Think showed intermediate latency, and ChatGPT-5.4 Thinking had the longest response times. Readability analyses revealed significant inter-model differences across all indices, with Gemini 3 Deep Think generating the most complex outputs and DeepSeek-V3.2 producing comparatively more readable responses. Clinical Risk Classification showed that most responses were categorized as having no clinically relevant risk or low clinical risk, while moderate-risk classifications were uncommon and no high-risk responses were identified. Under single-prompt, single-session benchmark conditions, potentially clinically relevant differences were observed among the models in MRONJ-related response quality, readability, response time, and clinical risk classification. DeepSeek-V3.2 and Gemini 3 Deep Think achieved higher expert-rated quality scores than ChatGPT-5.4 Thinking, while DeepSeek-V3.2 also showed advantages in speed and readability. These findings do not indicate general model superiority and support using LLMs only as adjunctive informational tools, not substitutes for specialist and guideline-based clinical decision-making.

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment