Regular expression-based learning to extract bodyweight values from clinical notes

Regular expression-based learning to extract bodyweight values from clinical notes
复制标题

DOI:
10.1016/j.jbi.2015.02.009
复制
发表时间:
2015-04-01
影响因子:
4.5
通讯作者:
Zeng-Treitler, Qing
Zeng-Treitler, Qing
中科院分区:
医学3区
文献类型:
--
作者:
Murtaugh, Maureen A.;Gibson, Bryan Smith;Zeng-Treitler, Qing

文献摘要

被引文献

相似文献

背景资料:体重相关测量(体重、身高、BMI、腹围)对于临床护理、研究和质量改进极其重要。这些和其他生命体征数据经常从电子健康记录的结构化表格中丢失。然而,它们通常被记录为临床笔记中的文本。在这个项目中,我们试图开发和验证一个学习算法,将提取体重相关的措施,从退伍军人管理局(VA)电子健康记录的临床笔记,以补充临床research.Methods中使用的结构化数据:我们开发了正则表达式发现提取器(REDEx),监督学习算法,从训练集生成正则表达式。REDEx生成的正则表达式,然后用于提取感兴趣的数值。方法:为了训练算法,我们创建了一个语料库的268个门诊初级保健笔记,由两个注释。该注释用于开发注释过程并识别与体重相关测量相关的术语,用于训练监督学习算法。另外300名门诊初级保健笔记的片段随后由两名评审员独立注释,以完成训练集。Inter-annotator agreements calculated.Methods:REDEx被应用到一个单独的测试集的3561笔记,以生成一个数据集的权重从文本中提取。我们估计的独特的个人,否则不会有体重相关的措施记录在CDW和额外的体重相关的措施,将另外capture.Results:REDEx的性能是:准确率= 98.3%,精确率= 98.8%,召回率= 98.3%,F = 98.5%。在来自3561个注释的体重数据集中,7.7%的注释包含无法作为结构化数据提供的体重相关测量。此外,2个额外的体重相关的措施,确定每一个人每年。结论:体重相关的措施,经常存储在临床笔记的文本。可以使用监督学习算法来提取该数据。临床护理,流行病学和质量改进工作的影响进行了讨论。(C)2015爱思唯尔公司All rights reserved.
Background: Bodyweight related measures (weight, height, BMI, abdominal circumference) are extremely important for clinical care, research and quality improvement. These and other vitals signs data are frequently missing from structured tables of electronic health records. However they are often recorded as text within clinical notes. In this project we sought to develop and validate a learning algorithm that would extract bodyweight related measures from clinical notes in the Veterans Administration (VA) Electronic Health Record to complement the structured data used in clinical research.Methods: We developed the Regular Expression Discovery Extractor (REDEx), a supervised learning algorithm that generates regular expressions from a training set. The regular expressions generated by REDEx were then used to extract the numerical values of interest.Methods: To train the algorithm we created a corpus of 268 outpatient primary care notes that were annotated by two annotators. This annotation served to develop the annotation process and identify terms associated with bodyweight related measures for training the supervised learning algorithm. Snippets from an additional 300 outpatient primary care notes were subsequently annotated independently by two reviewers to complete the training set. Inter-annotator agreement was calculated.Methods: REDEx was applied to a separate test set of 3561 notes to generate a dataset of weights extracted from text. We estimated the number of unique individuals who would otherwise not have bodyweight related measures recorded in the CDW and the number of additional bodyweight related measures that would be additionally captured.Results: REDEx's performance was: accuracy = 983%, precision = 98.8%, recall = 98.3%, F = 98.5%. In the dataset of weights from 3561 notes, 7.7% of notes contained bodyweight related measures that were not available as structured data. In addition 2 additional bodyweight related measures were identified per individual per year.Conclusion: Bodyweight related measures are frequently stored as text in clinical notes. A supervised learning algorithm can be used to extract this data. Implications for clinical care, epidemiology, and quality improvement efforts are discussed. (C) 2015 Elsevier Inc. All rights reserved.