Mining peripheral arterial disease cases from narrative clinical notes using natural language processing.

Mining peripheral arterial disease cases from narrative clinical notes using natural language processing.
复制标题

DOI:
10.1016/j.jvs.2016.11.031
复制
发表时间:
2017-06
影响因子:
4.3
通讯作者:
Arruda-Olson AM
Arruda-Olson AM
中科院分区:
医学2区
文献类型:
--
作者:
Afzal N;Sohn S;Abram S;Scott CG;Chaudhry R;Liu H;Kullo IJ;Arruda-Olson AM

文献摘要

被引文献

相似文献

下肢外周动脉疾病(PAD)是一种非常普遍的疾病,影响着全世界数百万人。我们开发了一种自然语言处理(NLP)系统,用于从临床叙述笔记中自动确定PAD病例,并将NLP算法的性能与计费代码算法进行了比较,使用踝臂指数(ABI)测试结果作为金标准。我们将NLP算法的性能与1)金标准ABI的结果; 2)基于相关ICD-9诊断代码(简单模型)的先前验证算法和3)ICD-9代码与程序代码(完整模型)的组合进行了比较。将1,569名PAD患者和对照的数据集随机分为训练(n= 935)和测试(n= 634)子集。我们在训练集中迭代地改进了NLP算法,包括叙述性笔记部分,笔记类型和服务类型,以最大限度地提高其准确性。在测试数据集上,与简单模型和全模型相比,NLP算法具有更好的精度(NLP:91.8%,完整模型:81.8%,简单模型:83%,P<.001),PPV(NLP:92.9%,全模型:74.3%,简单模型:79.9%,P<.001)和特异性(NLP:92.5%,全模型:64.2%,简单模型:75.9%,P<.001)。一个知识驱动的NLP算法,从临床笔记PAD病例的自动确定有更大的准确性比计费代码算法。我们的研究结果强调了NLP工具的潜力,可以快速有效地从电子健康记录中确定PAD病例,以促进临床调查,并最终通过临床决策支持来改善护理。
Lower extremity peripheral arterial disease (PAD) is highly prevalent and affects millions of individuals worldwide. We developed a natural language processing (NLP) system for automated ascertainment of PAD cases from clinical narrative notes and compared the performance of the NLP algorithm to billing code algorithms, using ankle-brachial index (ABI) test results as the gold standard. We compared the performance of the NLP algorithm to 1) results of gold standard ABI; 2) previously validated algorithms based on relevant ICD-9 diagnostic codes (simple model) and 3) a combination of ICD-9 codes with procedural codes (full model). A dataset of 1,569 PAD patients and controls was randomly divided into training (n= 935) and testing (n= 634) subsets. We iteratively refined the NLP algorithm in the training set including narrative note sections, note types and service types, to maximize its accuracy. In the testing dataset, when compared with both simple and full models, the NLP algorithm had better accuracy (NLP: 91.8%, full model: 81.8%, simple model: 83%, P<.001), PPV (NLP: 92.9%, full model: 74.3%, simple model: 79.9%, P<.001), and specificity (NLP: 92.5%, full model: 64.2%, simple model: 75.9%, P<.001). A knowledge-driven NLP algorithm for automatic ascertainment of PAD cases from clinical notes had greater accuracy than billing code algorithms. Our findings highlight the potential of NLP tools for rapid and efficient ascertainment of PAD cases from electronic health records to facilitate clinical investigation and eventually improve care by clinical decision support.