AI paper index

Rhythmic Monotonization in LLM-Generated Text Is Cross-Lingual: Evidence from Three Languages and Seven Models

2026-07-18 · Zenodo (CERN European Organization for Nuclear Research)

One-line summary

An AI research paper on Rhythmic Monotonization in LLM-Generated Text Is Cross-Lingual: Evidence from Three Languages and Seven Models.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

Our previous study of Japanese LLM-generated text found a two-layer fingerprint: vocabulary identifies which model wrote a text, while rhythm (suppressed sentence-length variation, uniform paragraph structure) shifts in the same direction in all models, like a shared machine accent. That study left an open question: is the accent a property of the machines, or an artifact of Japanese? We test the cross-lingual claim by replicating the rhythm protocol in English and Brazilian Portuguese: 7 LLMs (Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 4.8, GPT-3.5 Turbo, GPT-4o, GPT-OSS 20B, Llama 3.2 1B) x 10 themes x 5 trials per language, measured against 853 pre-ChatGPT Dev.to articles (EN) and 403 pre-ChatGPT TabNews posts (PT-BR), with the Japanese results of the prior study as the third point of comparison. Direction is universal: on all five core rhythm metrics (burstiness of sentence-length sequences in characters, words, and syllables; sentence-length CV in characters and words), every one of the 70 model x language cells shows the AI side more monotone than human (d < 0), before and after residualizing against document length. Magnitude is language-dependent but narrowly so: pooled character-burstiness effects are d = -0.96 (JA), -1.12 (EN), and -1.03 (PT), and the model ordering on burstiness is identical in all three languages (GPT-3.5 Turbo most monotone, GPT-OSS 20B closest to human). Comma density behaves in the opposite way: AI overshoots the human rate in English (d = +0.45) but undershoots it in Portuguese (d = -0.83), because AI writes commas at a nearly language-independent rate while human comma conventions differ twofold. Rhythm is a cross-lingual accent; punctuation is a language-specific dialect the models fail to acquire. A secondary result is generational: the current Claude tier shows roughly a third of GPT-3.5 Turbo's monotonization (d = -0.58 to -0.96 vs. -2.2 to -2.6), suggesting the accent is fading. Rhythm-based AI-text lint can reuse its metric directions across languages, but must calibrate thresholds per language.

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment