AI paper index

llms.txt across Romanian domains cited by AI engines

2026-08-26 · Zenodo (CERN European Organization for Nuclear Research)

One-line summary

An AI research paper on llms.txt across Romanian domains cited by AI engines.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

Original measurement of how many Romanian websites publish an llms.txt file, what is inside those files, and whether the same sites allow the crawlers that feed AI answers. Measured: 9 August 2026, one complete run. Frame: 87 Romanian-market domains that AI engines cite when answering Romanian-language questions, taken from the sources recorded in monitored answers. Inclusion rule: the domain appears among the 200 most-cited sources. Findings. 56 of 87 domains (64.4%) publish an llms.txt, making it already the majority practice in this niche rather than a curiosity. The 12 plugin-generated files average 25,070 bytes against 8,803 for the 44 hand-written ones, and carry neither the H1 nor the summary line the llmstxt.org specification asks for; only 39 files are spec-conformant and 8 contain no curated link. The single most-cited domain in the frame serves a plugin dump, which is evidence against llms.txt being what earns citations. Separately, the study separates answer-layer crawlers (OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-User, Claude-SearchBot) from training and archive scrapers (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider and others), following each vendor's current documentation: blocking the first removes a site from AI answers whatever its llms.txt says, blocking the second does not. Three domains block something; none of them block an answer-layer crawler. Limits. The study does not establish a causal relation between llms.txt and citations, and cannot separate its effect from domain age, authority or content. It does not measure whether the bots read the file, which would need server logs unavailable for third-party domains. A single run captures one day, which is why the scanner is published alongside the data. v1.1.0 (26 August 2026). Crawler classification corrected against each vendor's current documentation; GPTBot, ClaudeBot and Google-Extended are training-layer, not answer-layer. Domains blocking answer-engine crawlers falls from 2 to 0. The raw data is byte-identical to v1.0.0 — only the classification changed. Episode 4 of the Websem study series on AI-engine visibility. Study page: https://websem.ro/resurse/aeo/studiu-llms-txt-romania

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment