AI paper index

Knowledge Accuracy and Response Characteristics of an On-Device Large Language Model in Building Environmental Engineering: a Preliminary Case Study of EXAONE 4.0 1.2B Using ChatGPT-4o as a Cloud Reference

2026-08-24 · Architectural research

One-line summary

An AI research paper on Knowledge Accuracy and Response Characteristics of an On-Device Large Language Model in Building Environmental Engineering: a Preliminary Case Study of EXAONE 4.0 1.2B Using ChatGPT-4o as a Cloud Reference.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

This preliminary case study examines EXAONE 4.0 1.2B, an on-device large language model (LLM), in building environmental engineering, using ChatGPT-4o as a high-capability cloud reference rather than a size-matched competitor. Ten Korean-language items were evaluated: four multiple-choice, three single-step calculations, and three short-answer questions covering indoor environmental quality, energy, and post-occupancy evaluation. The same simple prompt condition was applied to both models, with no model-specific prompt optimization; therefore, the results represent one standardized condition rather than each model’s maximum performance. Both models answered all seven objective items correctly, although EXAONE showed terminology confusion in its reasoning on a thermal-environment item. Five doctoral-level experts rated the short-answer responses for accuracy, completeness, and logical consistency. Mean scores were 4.18 for ChatGPT-4o and 2.98 for EXAONE. Fleiss’ kappa was 0.020 and 0.044, while free-marginal kappa was 0.333 and 0.222, respectively; the latter values indicate fair, not moderate, agreement, so absolute expert scores require cautious interpretation. Qualitatively, ChatGPT-4o more consistently organized responses around concepts and scope, whereas EXAONE tended to provide specific technical, operational, and Korean-regulatory content but occasionally omitted canonical elements or added unsupported detail. Because the item pool is small and only one on-device model was tested, the findings are exploratory and cannot be generalized to building environmental engineering as a whole or to on-device LLMs as a class. The results support further investigation of on-device LLMs as supervised offline assistants where connectivity or external data transmission is constrained, but not as unsupervised tools for engineering calculation or compliance decisions.

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment