AI paper index

Evaluation of the accuracy and consistency of DeepSeek and ChatGPT in addressing type 1 diabetes mellitus-related queries in adults

2026-07-21 · Scientific Reports

One-line summary

An AI research paper on Evaluation of the accuracy and consistency of DeepSeek and ChatGPT in addressing type 1 diabetes mellitus-related queries in adults.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

Abstract Adult type 1 diabetes mellitus (T1DM) involves complex diagnosis, treatment, and long-term self-management, creating a need for accurate and accessible health education. Large language models (LLMs) are increasingly used for medical information seeking, yet their accuracy and consistency in adult T1DM-related queries remain insufficiently evaluated. A guideline-based comparative evaluation assessed DeepSeek-V3.2 and ChatGPT-5.0 using 22 English-language prompts derived from the 2021 ADA/EASD consensus report, covering basic knowledge, diagnosis and differential diagnosis, treatment, and complications. The prompts were submitted to both models twice, two weeks apart. Responses were independently evaluated by two blinded endocrinology specialists using a predefined four-point scoring rubric, with disagreements adjudicated by a third senior endocrinologist. Short-term consistency was assessed by expert judgment and TF-IDF cosine similarity. Inter-rater agreement was good (Cohen’s κ = 0.71). Expert-judged consistency was 95.45% (21/22) for both models; TF-IDF cosine similarity was 0.52 ± 0.09 for DeepSeek and 0.54 ± 0.09 for ChatGPT. Overall accuracy scores were 3.59 ± 0.59 and 3.77 ± 0.43, respectively, with no statistically significant difference (p = 0.102). Comprehensive ratings accounted for 63.64% and 77.27%, respectively, and mixed correct and incorrect or outdated information accounted for 4.55% and 0.00%. Both models may support adult T1DM-related health education, but outputs should be interpreted as supplementary educational material under professional guidance rather than as diagnostic or therapeutic advice.

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment