AI paper index

Beyond Mere Words: Tracing Large-Language-Model Style in Korean Academic Writing through Excess Vocabulary. A study of 316,538 Korean journal abstracts, 2018–2026, with a Vietnamese comparison on 47,165 abstracts

2026-08-26 · Zenodo (CERN European Organization for Nuclear Research)

One-line summary

An AI research paper on Beyond Mere Words: Tracing Large-Language-Model Style in Korean Academic Writing through Excess Vocabulary. A study of 316,538 Korean journal abstracts, 2018–2026, with a Vietnamese comparison on 47,165 abstracts.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

Excess vocabulary, a word's frequency above its pre-2023 trend, has measured how large language models changed scholarly English. We adapt it to Korean and Vietnamese with morphological units: 316,538 Korean abstracts from 2,258 KCI journals (2018 to August 2026) and, as an external comparison, 47,165 Vietnamese abstracts. Placebo runs put the false-positive level at 0.1–2.3 points. Korean abstracts show nothing in 2023, an onset in late 2024 and a rise through 2025 flattening in mid-2026: 시사하다 "suggest" appears in 21.4% of 2026 abstracts against 5.2% expected; plain verbs like 알아보다 "look into" fall to a quarter of trend. The single-word lower bound on LLM-processed abstracts (excess ratio at least 1.5) is 3.7% in 2024, 9.4% in 2025 and 16.2% in 2026; a split-half set bound is 7.5%, 19.3% and 31.1%. Control abstracts from three providers reproduce the rising words and sort markers by model generation. Model drafting puts the share at 69–98% (propagated ranges from 51%); rewriting by two API models reproduces the two headline words (66–90%) but not the broader vocabulary; polishing cannot produce it. In the English abstracts of the same articles the excess appears in 2023, a year earlier; articles whose English abstract carries the period's markers are 2.4 times as likely, in odds, to carry the Korean ones, but the Korean shift is present, at 68% of the rate, where the English side carries none. Vietnamese shows the same words a year later, too noisy for a bound. Code, tables and a dictionary are released. Working paper, version 5 (26 August 2026). Version 5 adds an article-level comparison with the English abstracts of the same KCI articles (Section 5.12, Tables 14 and 15, Figure 8) and rewrites the comparison with Koo, Kim and Kim (2026) from its full text; the version history is in Appendix J. The attached reproducibility package contains the analysis code, the per-year document-frequency tables for Korean and Vietnamese, the 1,180 generated positive-control abstracts (thirteen conditions from OpenAI, Anthropic and EXAONE models) and the manuscript consistency gate. A public Korean AI-style dictionary and checker built from the same measurements are at https://os.intframe.com/report/ai-style-dictionary-ko and https://os.intframe.com/report/ai-style-check-ko. Erratum (26 August 2026): the version 5 PDF carries three passages twice (the contribution item "Within-article comparison", the Section 3.1 paragraph on the English abstracts and the penultimate sentence of the Conclusion), and its Data availability section gives an outdated count of control abstracts (920; the count stated in the Methods, 1,180, is correct) and no DOI. A corrected copy, version 5.1, is at https://os.intframe.com/report/ai-style-lexicon (PDF: https://os.intframe.com/report/ai-style-lexicon-v5-1.pdf); no number, table, figure or claim changes. The corrections are carried into version 6.

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment