AI paper index
PHILIA TextAnalyzer CH7-v12 — Observational Pre-registration: Wikipedia Three-Segment Evidence Collection and LLM Reasoning Observation (106th DOI)
One-line summary
An AI research paper on PHILIA TextAnalyzer CH7-v12 — Observational Pre-registration: Wikipedia Three-Segment Evidence Collection and LLM Reasoning Observation (106th DOI).
Engineering notes
Engineering notes will be added by the aipentium editorial team.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。
Original abstract
This document is the pre-registration for the 106th DOI of the PHILIA OS research series (Trinity AI Research Team, Konyang University). It registers an observational study (NOT confirmatory) investigating whether Wikipedia three-segment evidence collection (front / middle / rear) improves delivery of claim-supporting sentences to a local LLM (qwen2.5:14b) compared to the v6 baseline (A arm accuracy 70.2%). Single experimental variable: The full Wikipedia Plain Text is divided into three equal segments by sentence count. The identical keyword-window method (collect_keyword_sentences, max_windows=1) is applied independently to each segment, yielding a maximum of 3 sentences per segment (9 sentences total, front→middle→rear fixed order, re-ranking prohibited). This isolates the positional coverage effect only, holding evidence volume constant relative to v6. Infrastructure changes (not experimental variables): MAX_EVIDENCE_CHARS=9,000 (empirical basis: v6 A arm max 1,884 chars × 3 segments = 5,652 chars; 9,000 includes margin); num_ctx=32,768. Internal activation observation track (parallel, observation-only): Logit Lens Entropy at L16/L24/L32/L40 via Qwen/Qwen2.5-14B-Instruct (fp16, RTX 4080 SUPER 16GB). No judgment intervention — Ollama (Q4) verdict authority maintained. Q4↔fp16 agreement rate recorded (O7). Pre-locked observation items (O1–O7): A arm accuracy vs v6 (O1); item-level improvement vs regression with item_id (O2); summary quality change (O3); logprob trajectory change per arm (O4); B/C/D arm stability (O5); internal activation values L16/L24/L32/L40 + logprob contrast (O6); Q4↔fp16 agreement rate (O7). Five forking paths closed: Step 0 dynamic injection; TF-IDF document selection (v11 confirmed net −3); 4-source parallel; transformers as judgment intervention; confirmatory endpoint declaration. Prior DOI: 105th Result — NOT_ESTABLISHED (10.5281/zenodo.21754299)Gate-0 adversarial review: Opus (Claude Opus) — Design PASS (v0.8, 2026-08-08); Code-level PASS (v12 script, 2026-08-09). Pandora (ChatGPT) — Code-level PASS (2026-08-09). Measurement is life. Description, not proof. Pre-registration protects us. — PHILIA OS | 0∞1∞0.5∞
Links and sources
Need this topic turned into a technical roadmap?
aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.
Request B2B AI research
Comments