AI paper index
Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting
One-line summary
An AI research paper on Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting.
Engineering notes
Engineering notes will be added by the aipentium editorial team.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。
Original abstract
Objective: Large language models (LLMs) are increasingly being used to develop medical education materials; however, it remains unclear how reliable, readable, or guideline-compliant the content generated by these models is for non-English-speaking patient groups. We evaluated the quality of Turkish melanoma patient education texts generated by seven frontier LLMs. Methods: A standardized 22-item Turkish prompt, built from international melanoma guidelines, was put to seven models in zero-shot sessions: ChatGPT 4.0 Turbo, Gemini 2.0 Flash, Claude 3.7 Sonnet, Grok 3, Qwen 2.5 Plus, DeepSeek R1, and Mistral Large 2. Each output was rated for readability (Ateşman Index), understandability, how clearly medical terminology was explained, scientific reliability (DISCERN instrument), empathy, and adherence to a 31-item guideline-based checklist. Model comparisons were summarized descriptively, using model-level absolute scores, score ranges, and rankings. Results: Model performance differed across readability, understandability, reliability, empathy, and guideline-adherence domains. DeepSeek R1 led on both readability (81.6) and understandability (23.5/25). Guideline adherence was strongest for Grok 3 and DeepSeek R1, at 96.8% and 93.5%, respectively, and Grok 3, DeepSeek R1, and Gemini 2.0 Flash each scored above 90% on the normalized total DISCERN measure. DeepSeek R1 also recorded the highest empathy score (90%). Gemini 2.0 Flash had the lowest readability score and produced the longest output (Ateşman 65.8). None of the models provided citations or verifiable sources, so every model received the lowest possible DISCERN Source Reliability score; Mistral Large 2 showed the weakest overall performance. Conclusion: How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next. A few produced text that was clear, empathetic, and reasonably guideline-concordant, but the lack of verifiable citations and uneven guideline coverage remain genuine limitations. These findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.
Links and sources
Need this topic turned into a technical roadmap?
aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.
Request B2B AI research
Comments