AI paper index
Rhythmic Monotonization in LLM-Generated Text Is Cross-Lingual: Evidence from Three Languages and Seven Models
One-line summary
An AI research paper on Rhythmic Monotonization in LLM-Generated Text Is Cross-Lingual: Evidence from Three Languages and Seven Models.
Engineering notes
Engineering notes will be added by the aipentium editorial team.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。
Original abstract
Our previous study of Japanese LLM-generated text found a two-layer fingerprint: vocabulary identifies which model wrote a text, while rhythm (suppressed sentence-length variation, uniform paragraph structure) shifts in the same direction in all models, like a shared machine accent. That study left an open question: is the accent a property of the machines, or an artifact of Japanese? We test the cross-lingual claim by replicating the rhythm protocol in English and Brazilian Portuguese: 7 LLMs (Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 4.8, GPT-3.5 Turbo, GPT-4o, GPT-OSS 20B, Llama 3.2 1B) x 10 themes x 5 trials per language, measured against 853 pre-ChatGPT Dev.to articles (EN) and 403 pre-ChatGPT TabNews posts (PT-BR), with the Japanese results of the prior study as the third point of comparison. Direction is universal: on all five core rhythm metrics (burstiness of sentence-length sequences in characters, words, and syllables; sentence-length CV in characters and words), every one of the 70 model x language cells shows the AI side more monotone than human (d < 0), before and after residualizing against document length. Magnitude is language-dependent but narrowly so: pooled character-burstiness effects are d = -0.96 (JA), -1.12 (EN), and -1.03 (PT), and the model ordering on burstiness is identical in all three languages (GPT-3.5 Turbo most monotone, GPT-OSS 20B closest to human). Comma density behaves in the opposite way: AI overshoots the human rate in English (d = +0.45) but undershoots it in Portuguese (d = -0.83), because AI writes commas at a nearly language-independent rate while human comma conventions differ twofold. Rhythm is a cross-lingual accent; punctuation is a language-specific dialect the models fail to acquire. A secondary result is generational: the current Claude tier shows roughly a third of GPT-3.5 Turbo's monotonization (d = -0.58 to -0.96 vs. -2.2 to -2.6), suggesting the accent is fading. Rhythm-based AI-text lint can reuse its metric directions across languages, but must calibrate thresholds per language.
Links and sources
Need this topic turned into a technical roadmap?
aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.
Request B2B AI research
Comments