AI paper index

EHTEST – A systematically sampled and annotated corpus for the study of human–nature relations in Estonian thought and culture

2026-08-23 · Zenodo (CERN European Organization for Nuclear Research)

One-line summary

An AI research paper on EHTEST – A systematically sampled and annotated corpus for the study of human–nature relations in Estonian thought and culture.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

EHTEST (Environmental Humanities Texts in ESTonian) is a representative corpus of Estonian-language non-fiction texts (‘essays’ sensu lato) published since the late 19th century. The corpus is designed to enable systematic comparison of how human–nature relationships have been conceptualized and described across historical periods and intellectual fields, and through different rhetorical and representational strategies in Estonian non-fiction writing. The principal criterion for inclusion was that a text addresses relationships between humans (or human societies) and nature, including interactions between the two systems. Representativeness is pursued through two complementary sets: (i) the main corpus – based on stratified sampling of authors from three groups (literature, science, and other fields), with each author represented by three texts; where more texts were available, the three were selected to reflect diversity in the author’s thought; and (ii) the supplementary corpus – individual texts by authors not represented in the main corpus, selected to compensate for temporal biases inherent in the main sampling design and to provide additional coverage of significant authors, themes and modes of representation. Both sets remain open to further expansion. The current version, EHTEST200, contains 180 texts in the main corpus (three groups × 20 authors × three texts) and 20 texts by 20 additional authors in the supplementary corpus from 1878-2025. Following predefined protocols, digitised Estonian source texts were summarised and indexed in English and supplemented with short author biographies using ChatGPT. The resulting material was subsequently checked against the original texts and subjected to terminological and methodological harmonisation. The deposited files comprise: (1) metadata (.xlsx), including text IDs, publication details and comments, and three sets of keywords representing natural-science, social-science and humanities content; (2) Methods description (.docx); and (3)–(4) English author biographies and standardised text summaries separately for the main and supplementary corpus. Revisions. Some typographic errors in the original metadata table have been corrected in EHTEST200_v2. Citation: When referring to or quoting an original Estonian text included in EHTEST, please cite its original publication or another appropriate source edition (e.g., as specified in the corpus metadata). When using EHTEST as a systematically compiled text corpus or its English-language data and annotations, please cite the relevant version of the EHTEST dataset. When referring to the conceptual and methodological approach underlying the corpus, please cite: Lõhmus, A. 2025. Making sense of human–nature relationships in essays: A functional view on the Estonian essay [Inimese ja looduse suhte mõtestamine esseistikas: funktsionaalne vaade eesti esseistikale]. Keel ja Kirjandus 68(4): 311–337. https://doi.org/10.54013/kk808a3

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment