AI paper index
A study on the reliability of ChatGPT-based English writing assessment: Internal and external perspectives
One-line summary
An AI research paper on A study on the reliability of ChatGPT-based English writing assessment: Internal and external perspectives.
Engineering notes
Engineering notes will be added by the aipentium editorial team.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。
Original abstract
With the advancement of big data and artificial intelligence, natural language processing (NLP) has been increasingly integrated into educational assessment, facilitating a shift from human to automated scoring in English writing assessment. This study investigates how prompt design, guided by the TELeR taxonomy (Santu & Feng, 2023), and sample training affects ChatGPT-4's scoring reliability. A sample of 120 IELTS-style essays was scored across four analytic dimensions--Task Response (TR), Coherence and Cohesion (CC), Lexical Resource (LR), and Grammatical Range and Accuracy (GRA)--as well as holistically. Internal (intra-rater) reliability was measured via ICC(3,1) across three scoring rounds, while external (inter-rater) reliability was assessed by integrating ChatGPT-4 as an additional rater into a panel of nine human scorers using ICC(2,1). The results revealed that ChatGPT-4's internal consistency was moderate and was markedly improved by detailed prompts on content-related dimensions (TR/CC), but showed minimal gains on linguistic dimensions (LR/GRA). Training with human-scored exemplars under the better prompt reduced discrepancies with human ratings for external reliability, yet slightly decreased internal consistency for most dimensions. These findings position TELeR as a framework linking prompt engineering to psychometric theory and underscore the value of dimension-specific prompt design for reliable AI-based writing assessment.
Links and sources
Need this topic turned into a technical roadmap?
aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.
Request B2B AI research
Comments