AI paper index

E091 - Recognition vs Discovery: Owned-Namespace Citation Ownership (Pre-Registration, T0)

2026-07-22 · Zenodo (CERN European Organization for Nuclear Research)

One-line summary

An AI research paper on E091 - Recognition vs Discovery: Owned-Namespace Citation Ownership (Pre-Registration, T0).

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

E091 — Recognition vs Discovery: Owned-Namespace Citation Ownership. Pre-Registration (T0). Registry: The GEO Lab Experiments DB, Experiment ID E091. (Note on identifiers: an earlier internal calendar entry carried the stale label "E088"; E088 is a separate registry record — the E045 infrastructure-vs-citation cross-section. E091 is the canonical ID for this study. Stated here to pre-empt any bookkeeping ambiguity.)Investigator: Artur Ferreira, The GEO Lab (thegeolab.net) · ORCID 0009-0004-4072-9741Validity path: Path B — Observational, two-arm within-term comparison. Tier: T1 proprietary.Study window: T0 collection 22–26 July 2026. Query-only; no site changes; not gated by the E059 hub freeze.Deposit timing: This document is deposited BEFORE any data collection. No query in the registered grid has been executed at deposit time. 1. MOTIVATION Sharma (Discovery Gap preprint, Martinez survey corpus) reports that ChatGPT recognises 99.4% of 112 startups when named but surfaces them in only 3.32% of organic category-discovery queries; Perplexity 94.3% vs 8.29%. The discovery-vs-recognition cell contains exactly one study, zero peer-reviewed. The GEO Lab's proven finding F-073-C (coined-terminology citation ownership) is currently unscoped against this distinction. E091 scopes it with the Lab's own data. 2. HYPOTHESES (registered verbatim) H1: For Lab-owned coined terms, citation/mention rate under named-entity queries exceeds rate under category-intent queries by more than the 22pp noise floor (E016). H0-alt worth stating: if category-intent rate exceeds noise floor, the recognition-gate model is wrong and F-073-C generalises further than Sharma implies — either outcome publishes. 3. DESIGN Two-arm within-term comparison across six terms (three Lab-owned coined terms, three non-owned controls), two engines, five query stems per arm per term = 60 stems. Arms. Arm A (recognition): queries that name the term. Arm B (discovery): category-intent queries for which the term's canonical page would be a correct answer but the term is not named. Arm B stems were drafted before any data collection (blind to results by construction), are fixed at this deposit, and admit no post-hoc additions. Terms and canonical pages: OWNED-1 — GEO Stack — owned — https://thegeolab.net/geo-stack/ OWNED-2 — Entity Reinforcement — owned — https://thegeolab.net/entity-reinforcement/ OWNED-3 — Brand Citation Index (BCI) — owned — https://thegeolab.net/geo-brand-citation-index/ CONTROL-1 — Generative Engine Optimization — control — https://thegeolab.net/what-is-generative-engine-optimisation/ CONTROL-2 — AI Overviews — control — no canonical Lab page (related coverage: https://thegeolab.net/platform-specific-geo/) CONTROL-3 — llms.txt — control — https://thegeolab.net/llms-txt-what-it-does/ Controls are established terms in the same topic space where thegeolab.net has published content but did not coin the term. Control canonical coverage is deliberately uneven and is not a defect: CONTROL-1 and CONTROL-3 have dedicated Lab pages, whereas CONTROL-2 (AI Overviews) has none — the Lab has never published a definitional page on that feature. This is precisely the condition a control is meant to instantiate (a term in the topic space that the Lab neither coined nor owns), and it is registered here rather than remedied, since publishing content on a control term inside the study window would itself be an intervention. Term-selection note, disclosed: the owned slot OWNED-2 was finalised at pre-registration as Entity Reinforcement, replacing an earlier candidate (Surface-Layer Principle) that has no public canonical page and therefore fails both the Arm B answerability requirement and the study's own Stage 0 eligibility logic. The replacement was made before any data collection. Site-change hold. Because llms.txt is a control term, the planned deployment of the site's llms-full.txt file is held until after 26 July 2026. No site change touching any registered term's footprint will be made during the T0 window. 4. INSTRUMENT ChatGPT and Perplexity arms — T0 checks are performed manually on each engine's consumer web interface in a logged-out Chrome incognito session, one fresh browser profile per stem per repetition (unique wiped user-data directory, incognito mode), one query per session. Repetitions. Three parallel repetitions per stem per platform, executed within a single collection sitting (repetitions are parallel fresh-profile sessions, not time-separated runs). Pre-committed sequential-precision extension (per Schulte/Martinez §6.2, fixed here so it is not optional stopping): when a stem's three repetitions are non-unanimous on the citation outcome, that stem × platform is extended to five repetitions in the same sitting. Instrument selection, disclosed. Two automated alternatives were available and deliberately not used for the confirmatory grid: (a) the DataForSEO ChatGPT endpoint, and (b) the OpenAI API (Chat Completions or Responses-with-web_search). Both were excluded on construct-validity grounds — a bare API completion performs no retrieval and measures parametric recall rather than product citation behaviour, and any retrieval-enabled API surface uses a different retrieval/ranking stack and system prompt than the consumer products this study targets. This is an a priori design decision, not a post-hoc substitution. Known limitation: consumer surfaces expose no pinnable model string, so the underlying model version cannot be frozen and may drift within the T0 window; the grid completes by 26 July 2026 to bound this drift. Collection schedule. 12 stems/day across five days (22–26 July), three sittings per day, each sitting pairing one owned and one control term across both arms so that collection day and time-of-day are balanced against the owned-vs-control contrast. Term-to-sitting assignment is fixed across days; consequently within-owned term comparisons carry a time-of-day component and are reported as exploratory only. Launcher scripts and their launch logs are archived with this deposit (hashes recorded); they constitute the confirmatory collection instrument. 5. CODED FIELDS (per session) mode — 1 = substantive answer returned (including a disambiguation response); 0 = refusal / empty / error. Mode-0 sessions are coded explicitly and REMAIN IN THE DENOMINATOR. A substantive answer with no sources is mode = 1. retrieval_shown — 1 = any sources displayed with the answer (Perplexity sources panel present, or ChatGPT inline citations present); 0 = none. Reported separately as a retrieval proxy; the citation outcome is interpretable only where retrieval_shown = 1. citation (primary DV) — 1 = thegeolab.net present in the engine's shown sources panel (Perplexity's sources list; ChatGPT's inline citation set), assessed at REGISTRABLE-DOMAIN level after URL canonicalisation, deduplicated by canonical URL. The sources-panel definition is registered here explicitly; it is broader than an inline-footnoted-only definition and is not comparable to prior experiments coded on the narrower rule. mention (secondary DV) — 1 = the brand or term is named in the answer prose, excluding text appearing only within citation markers, source labels, or URLs (so citation-format differences between engines cannot inflate the mention rate). canonical_url — canonicalisation strips protocol variants, www, trailing slash, fragment, and all query/tracking parameters (including utm_source=chatgpt.com appended by ChatGPT), lowercases the host, and resolves to the registrable domain for the DV. session_fresh — Y required for scoring; sessions flagged N are re-run. Coding provenance, disclosed. Query stems were AI-assisted drafts reviewed, edited, and approved by the investigator before this deposit. First-pass outcome coding may be AI-assisted from raw response transcripts; an independent HUMAN second coder re-scores a random ≥20% sample on all four fields, with Cohen's κ reported per field (convention: κ ≥ 0.60 minimum (hard halt below), ≥ 0.65 target — the Lab's standing inter-coder gate, applied unchanged from E085). The independent human κ pass carries the reliability claim. 6. REGISTERED QUERY STEMS (frozen at deposit; executed verbatim, no additions) OWNED-1 · GEO Stack — Arm A (S01–S05)S01. What is the GEO Stack?S02. Explain the five layers of the GEO Stack framework.S03. Who developed the GEO Stack?S04. How do I use the GEO Stack to audit my content?S05. How is the GEO Stack different from traditional SEO? OWNED-1 · GEO Stack — Arm B (S06–S10)S06. Is there a layered framework for diagnosing why my content never gets cited in AI search answers?S07. Do I need to rank in Google before ChatGPT or Perplexity can cite my page?S08. In what order should I fix problems when my pages rank well but AI answers never cite them?S09. Why doesn't adding schema markup help a page that AI systems never retrieve in the first place?S10. Is there a measurement-first framework for AI search visibility backed by controlled experiments? OWNED-2 · Entity Reinforcement — Arm A (S11–S15)S11. What is Entity Reinforcement?S12. What is entity gravity and how does Entity Reinforcement build it?S13. Who coined the term Entity Reinforcement?S14. How do I apply Entity Reinforcement to my website content?S15. What is the difference between Entity Reinforcement and keyword optimisation? OWNED-2 · Entity Reinforcement — Arm B (S16–S20)S16. Why does my website get cited by AI search engines for a query one day but not the next, even though nothing changed?S17. Does using different names for the same concept across my site hurt my chances of being cited in AI answers?S18. How do AI systems decide which topics a website is an authority on?S19. How can I test whether AI search engines have a stable association between my site and my topic?S20. What should internal link anchor text say to help AI systems understand what my pages are about? OWNED-3 · Brand Citati

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment