AI paper index

PHILIA TextAnalyzer CH7-v12 — Observational Pre-registration: Wikipedia Three-Segment Evidence Collection and LLM Reasoning Observation (106th DOI)

2026-08-09 · Zenodo (CERN European Organization for Nuclear Research)

One-line summary

An AI research paper on PHILIA TextAnalyzer CH7-v12 — Observational Pre-registration: Wikipedia Three-Segment Evidence Collection and LLM Reasoning Observation (106th DOI).

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

This document is the pre-registration for the 106th DOI of the PHILIA OS research series (Trinity AI Research Team, Konyang University). It registers an observational study (NOT confirmatory) investigating whether Wikipedia three-segment evidence collection (front / middle / rear) improves delivery of claim-supporting sentences to a local LLM (qwen2.5:14b) compared to the v6 baseline (A arm accuracy 70.2%). Single experimental variable: The full Wikipedia Plain Text is divided into three equal segments by sentence count. The identical keyword-window method (collect_keyword_sentences, max_windows=1) is applied independently to each segment, yielding a maximum of 3 sentences per segment (9 sentences total, front→middle→rear fixed order, re-ranking prohibited). This isolates the positional coverage effect only, holding evidence volume constant relative to v6. Infrastructure changes (not experimental variables): MAX_EVIDENCE_CHARS=9,000 (empirical basis: v6 A arm max 1,884 chars × 3 segments = 5,652 chars; 9,000 includes margin); num_ctx=32,768. Internal activation observation track (parallel, observation-only): Logit Lens Entropy at L16/L24/L32/L40 via Qwen/Qwen2.5-14B-Instruct (fp16, RTX 4080 SUPER 16GB). No judgment intervention — Ollama (Q4) verdict authority maintained. Q4↔fp16 agreement rate recorded (O7). Pre-locked observation items (O1–O7): A arm accuracy vs v6 (O1); item-level improvement vs regression with item_id (O2); summary quality change (O3); logprob trajectory change per arm (O4); B/C/D arm stability (O5); internal activation values L16/L24/L32/L40 + logprob contrast (O6); Q4↔fp16 agreement rate (O7). Five forking paths closed: Step 0 dynamic injection; TF-IDF document selection (v11 confirmed net −3); 4-source parallel; transformers as judgment intervention; confirmatory endpoint declaration. Prior DOI: 105th Result — NOT_ESTABLISHED (10.5281/zenodo.21754299)Gate-0 adversarial review: Opus (Claude Opus) — Design PASS (v0.8, 2026-08-08); Code-level PASS (v12 script, 2026-08-09). Pandora (ChatGPT) — Code-level PASS (2026-08-09). Measurement is life. Description, not proof. Pre-registration protects us. — PHILIA OS | 0∞1∞0.5∞

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment