AI paper index

Blind spots in artificial intelligence systems: poor identification of self-generated medical images–evidence-based cross-sectional study

2026-08-11 · Frontiers in Medicine

One-line summary

An AI research paper on Blind spots in artificial intelligence systems: poor identification of self-generated medical images–evidence-based cross-sectional study.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

Objectives Artificial intelligence (AI) systems are increasingly adopted in scientific visualization, clinical sciences, medical education, and health sciences research. Despite its impressive generative capabilities, the AI role remains poorly established in some areas. This study aims to investigate whether AI systems, ChatGPT-4 and Google Gemini 1.5 Pro, can identify their own generated images. Methods In this study, AI systems, OpenAI’s ChatGPT-4 and Google Gemini 1.5 Pro, were employed to generate and identify medical science images. AI-generated images in medical science span 16 clinical illustrations: 8 anatomical and 8 pathophysiological mechanisms. The images were subsequently reintroduced into the same system for identification. The descriptive statistics, including numbers, percentages, and accuracy rates, across the images were analyzed. For a correct answer or an incorrect answer, a score of 1 was allocated; a p -value less than 0.05 was considered significant. Results The results revealed that Artificial Intelligence models, ChatGPT and Google Gemini, were significantly less able to identify AI-generated images correctly. There were significantly higher incorrect rates than correct identification rates for AI-generated images overall (27/32 incorrect vs. 5/32 correct; 84.37% vs. 15.62%; p = 0.001). ChatGPT showed an overall accuracy of 18.75% ( p = 0.021), while Google Gemini demonstrated an overall accuracy of 12.5% ( p = 0.004). The poorest performance was observed in Google Gemini anatomy images, with no correct identifications (0%, p = 0.008). Conclusion Artificial Intelligence models, ChatGPT and Google Gemini, have limited ability to identify AI-generated images correctly. The results demonstrate structured, architecture-dependent patterns of image identification dysfunction. The study emphasizes implications for AI governance and validation protocols in research, medicine, and medical education.

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment