AI paper index

Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus

2026-08-01 · Figshare

One-line summary

An AI research paper on Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

Datasets that ship automatically generated feature annotations invite a question that is rarely asked of them: would a human agree with those labels? This report answers that question for the Objective Projection corpus, a Turkish narrative dataset whose scenes carry a per-scene applied_rules field produced by a rule-based detector over six craft features — two prohibitions<br>(explicit emotion labelling, simile) and four positive techniques (materialized metaphor, microfocus, temporal anchor, atmosphere contradiction).<br><br>Three studies are reported. Study 1 (𝑛 = 120) scores the detector against blind labels produced by the scheme’s own author. Study 2 (𝑛 = 100, a disjoint scene set) scores the detector plus Gemini 2.5 Flash and Grok against an independent non-expert human rater whose labels were locked before any machine ran. Study 2b re-runs the identical protocol with Claude Fable 5<br>(High) and ChatGPT 5.5. The central result concerns one rule. On materialized metaphor — the feature closest to the<br>methodology’s theoretical core — the five machine labellers returned positive rates of 0, 1, 40, 72 and 78 out of 100 scenes, against a human count of 9. Cohen’s 𝜅 was at or indistinguishable from chance for five of the six machine labellers, across both human references, on both scene sets: 0.004, 0.015, 0.000, 0.019, 0.027. Raw agreement, by contrast, ranged from 74.7% to 84.5%, an artefact of class imbalance rather than a sign of competence.<br><br>We deliberately do not resolve the finding into a single story. Two readings survive the data: that the feature is genuinely inferential and beyond current automatic detection, or that the 1 rule’s definition is not yet operational enough for any rater to apply consistently — including the human. Distinguishing them requires a second independent human rater, which this report<br>does not have and therefore does not claim.<br><br>Keywords: inter-rater reliability, Cohen’s kappa, annotation quality, class imbalance, LLM-asannotator, computational narratology, show-don’t-tell, Turkish corpus, Bulut Doctrine

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment