AI paper index
Cutting chart-review time and improving database accuracy in inflammatory bowel disease with human-in-the-loop large language models
One-line summary
An AI research paper on Cutting chart-review time and improving database accuracy in inflammatory bowel disease with human-in-the-loop large language models.
Engineering notes
Engineering notes will be added by the aipentium editorial team.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。
Original abstract
Abstract Background Manual abstraction of electronic health records into research databases is a major bottleneck in clinical research, limiting scalability and introducing error. This challenge is particularly acute in inflammatory bowel disease surveillance, where clinically relevant variables are distributed across extensive longitudinal documentation. We evaluated whether a reproducible, open-source, human-in-the-loop workflow based on large language models could outperform standard manual chart review in both efficiency and accuracy. Methods We developed a locally deployed, source-linked extraction pipeline using open-source large language models to recreate an existing inflammatory bowel disease surveillance database. The system employed a two-stage, activation-based architecture that extracted structured variables from clinical notes and returned each value with supporting source citations. In a controlled user study, four clinicians each annotated 20 patients: 10 by manual chart review and 10 by reviewing and correcting model-generated outputs using a custom, source-aware user interface. Primary outcomes were extraction time per patient and per variable, and extraction accuracy relative to ground truth. Time outcomes were compared using two-sided Mann–Whitney U tests due to non-normal distributions, with effect sizes and confidence intervals reported. Accuracy differences were summarized using absolute risk differences with confidence intervals. Results Median extraction time per patient decreased from 9.4 min with manual review to 3.6 min with model assistance, yielding a typical time saving of 5.27 min per patient and a large effect size ( r = 0.747, 95% confidence interval 0.576–0.885). Extraction accuracy improved from 68% with manual abstraction to 89% with model-assisted annotation (risk difference 0.213; 95% confidence interval 0.102–0.316). Accuracy gains were greatest for variables requiring synthesis across multiple clinical notes, while performance on high-salience variables was comparable across workflows. Conclusions An open-source, human-verified workflow using large language models can accelerate electronic health record abstraction while improving accuracy. By releasing both the extraction pipeline and user interface software, this study provides a reproducible and deployable template for scalable clinical database curation and supports broader adoption of transparent artificial intelligence methods in clinical research.
Links and sources
Need this topic turned into a technical roadmap?
aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.
Request B2B AI research
Comments