AI paper index
Multimodal large language model–assisted workflow for image-dependent very short answer question generation and response grading in neurosurgical residency assessment: development and validation study
One-line summary
An AI research paper on Multimodal large language model–assisted workflow for image-dependent very short answer question generation and response grading in neurosurgical residency assessment: development and validation study.
Engineering notes
Engineering notes will be added by the aipentium editorial team.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。
Original abstract
Very short answer questions (VSAQs) reduce cueing but require substantial specialist effort to develop and grade, particularly when they involve neuroimaging. We evaluated an expert-supervised workflow using multimodal large language models (LLMs) to generate image-dependent neurosurgical VSAQs and support grading of resident responses. In this three-phase, single-center study, two multimodal LLMs generated one VSAQ from each of 36 deidentified neurosurgical cases using identical clinical information and three MRI images. Three senior neurosurgeons rated the 72 candidate items for scientific accuracy, image dependency, educational value, and clarity. The higher-scoring item for each case formed a 36-item Golden Set administered to 10 neurosurgery residents. Two senior neurosurgeons independently scored all 360 responses using a 0–2 rubric and resolved disagreements by consensus. ChatGPT graded the same responses using case-specific structured inputs and a fixed prompt. Between-model ratings were compared using paired t tests; agreement was assessed using exact agreement and quadratic weighted kappa, and participant-level totals using Pearson correlation. Both models received high ratings for scientific accuracy, educational value, and clarity, with no significant between-model differences. ChatGPT had higher image dependency ratings than Gemini (mean 4.06, SD 0.44 vs. 3.58, SD 0.52; mean difference 0.48, 95% CI 0.26–0.71; P < 0.001), whereas total scores did not differ significantly (mean 16.34, SD 0.86 vs. 15.93, SD 1.12; P = 0.115). Across 360 resident responses, 275 (76.4%) were fully correct, 28 (7.8%) partially correct, and 57 (15.8%) incorrect; total scores ranged from 47 to 63 out of 72. Before consensus, the experts agreed on 347 responses (96.4%; quadratic weighted kappa = 0.967). AI grading agreed exactly with the final expert scores for 324 responses (90.0%; quadratic weighted kappa = 0.905), and participant-level totals were strongly correlated (Pearson r = 0.970). Under the structured conditions evaluated, multimodal LLMs showed preliminary feasibility for expert-supervised drafting and first-pass grading of image-dependent neurosurgical VSAQs. Larger multicenter studies and external validation are needed before routine implementation or use in high-stakes assessment.
Links and sources
Need this topic turned into a technical roadmap?
aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.
Request B2B AI research
Comments