AI paper index
Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus
One-line summary
An AI research paper on Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus.
Engineering notes
Engineering notes will be added by the aipentium editorial team.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。
Original abstract
Datasets that ship automatically generated feature annotations invite a question that is rarely asked of them: would a human agree with those labels? This report answers that question for the Objective Projection corpus, a Turkish narrative dataset whose scenes carry a per-scene applied_rules field produced by a rule-based detector over six craft features — two prohibitions<br>(explicit emotion labelling, simile) and four positive techniques (materialized metaphor, microfocus, temporal anchor, atmosphere contradiction).<br><br>Three studies are reported. Study 1 (𝑛 = 120) scores the detector against blind labels produced by the scheme’s own author. Study 2 (𝑛 = 100, a disjoint scene set) scores the detector plus Gemini 2.5 Flash and Grok against an independent non-expert human rater whose labels were locked before any machine ran. Study 2b re-runs the identical protocol with Claude Fable 5<br>(High) and ChatGPT 5.5. The central result concerns one rule. On materialized metaphor — the feature closest to the<br>methodology’s theoretical core — the five machine labellers returned positive rates of 0, 1, 40, 72 and 78 out of 100 scenes, against a human count of 9. Cohen’s 𝜅 was at or indistinguishable from chance for five of the six machine labellers, across both human references, on both scene sets: 0.004, 0.015, 0.000, 0.019, 0.027. Raw agreement, by contrast, ranged from 74.7% to 84.5%, an artefact of class imbalance rather than a sign of competence.<br><br>We deliberately do not resolve the finding into a single story. Two readings survive the data: that the feature is genuinely inferential and beyond current automatic detection, or that the 1 rule’s definition is not yet operational enough for any rater to apply consistently — including the human. Distinguishing them requires a second independent human rater, which this report<br>does not have and therefore does not claim.<br><br>Keywords: inter-rater reliability, Cohen’s kappa, annotation quality, class imbalance, LLM-asannotator, computational narratology, show-don’t-tell, Turkish corpus, Bulut Doctrine
Links and sources
Need this topic turned into a technical roadmap?
aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.
Request B2B AI research
Comments