Automatic generation of case-detection algorithms to identify children with asthma from large electronic health record databases

Automatic generation of case-detection algorithms to identify children with asthma from large electronic health record databases
复制标题

DOI:
10.1002/pds.3438
复制
发表时间:
2013-08-01
影响因子:
2.6
通讯作者:
Schuemie, Martijn J.
Schuemie, Martijn J.
中科院分区:
医学4区
文献类型:
--
作者:
Afzal, Zubair;Engelkes, Marjolein;Schuemie, Martijn J.

文献摘要

被引文献

相似文献

大多数电子健康记录数据库包含非结构化的自由文本叙述,这是不容易分析。病例检测算法通常是手动创建的,并且通常仅依赖于使用编码信息,例如国际疾病分类第9版代码。我们应用机器学习的方法来生成和评估一个自动化的情况下,检测算法,使用自由文本和编码的信息来识别asthma cases.Methods的综合初级保健信息(IPCI)数据库中搜索潜在的哮喘患者年龄在5-18岁,使用广泛的查询哮喘相关的代码,药物和自由文本。通过将潜在患者手动注释为明确的、可能的或可疑的哮喘病例或非哮喘病例,创建了5032例患者的训练集。然后,使用规则学习程序RIPPER生成算法来区分案例和非案例。过采样方法被用来平衡自动算法的性能,以满足我们的研究要求。结果所选算法在仅识别明确哮喘病例时的阳性预测值(PPV)为0.66,敏感性为0.98,特异性为0.95,在识别明确和可能哮喘病例时的PPV为0.82,敏感性为0.96,特异性为0.90;和PPV为0.57,灵敏度为0.95,特异性为0.67的情况下确定明确的,可能的,可疑的哮喘cases.Conclusions自动算法显示出良好的性能,在检测哮喘病例利用自由文本和编码数据。该算法将促进IPCI数据库中哮喘的大规模研究。版权所有(C)2013约翰威利父子有限公司
Purpose Most electronic health record databases contain unstructured free-text narratives, which cannot be easily analyzed. Case-detection algorithms are usually created manually and often rely only on using coded information such as International Classification of Diseases version 9 codes. We applied a machine-learning approach to generate and evaluate an automated case-detection algorithm that uses both free-text and coded information to identify asthma cases.Methods The Integrated Primary Care Information (IPCI) database was searched for potential asthma patients aged 5-18 years using a broad query on asthma-related codes, drugs, and free text. A training set of 5032 patients was created by manually annotating the potential patients as definite, probable, or doubtful asthma cases or non-asthma cases. The rule-learning program RIPPER was then used to generate algorithms to distinguish cases from non-cases. An over-sampling method was used to balance the performance of the automated algorithm to meet our study requirements. Performance of the automated algorithm was evaluated against the manually annotated set.Results The selected algorithm yielded a positive predictive value (PPV) of 0.66, sensitivity of 0.98, and specificity of 0.95 when identifying only definite asthma cases; a PPV of 0.82, sensitivity of 0.96, and specificity of 0.90 when identifying both definite and probable asthma cases; and a PPV of 0.57, sensitivity of 0.95, and specificity of 0.67 for the scenario identifying definite, probable, and doubtful asthma cases.Conclusions The automated algorithm shows good performance in detecting cases of asthma utilizing both free-text and coded data. This algorithm will facilitate large-scale studies of asthma in the IPCI database. Copyright (C) 2013 John Wiley & Sons, Ltd.