Enhancing Comparative Effectiveness Research With Automated Pediatric Pneumonia Detection in a Multi-Institutional Clinical Repository: A PHIS+ Pilot Study.

Enhancing Comparative Effectiveness Research With Automated Pediatric Pneumonia Detection in a Multi-Institutional Clinical Repository: A PHIS+ Pilot Study.
复制标题

DOI:
10.2196/jmir.6887
复制
发表时间:
2017-05-15
影响因子:
7.4
通讯作者:
Shah S
Shah S
中科院分区:
医学2区
文献类型:
--
作者:
Meystre S;Gouripeddi R;Tieder J;Simmons J;Srivastava R;Shah S

文献摘要

被引文献

相似文献

社区获得性肺炎是儿科发病率的主要原因。行政数据经常用于进行具有足够样本量的比较有效性研究(CER),以加强对重要结果的检测。然而,由于放电诊断代码的准确性不同,这类研究容易出现误分类错误。这项研究的目的是开发一种自动化的、可扩展的、准确的方法来使用胸部成像报告来确定儿童肺炎的存在或不存在。开发了多机构PHIS+临床资料库,通过扩展具有详细临床数据的儿童医院管理数据库来支持儿科CER。为了开发一种可扩展的方法来更准确地找到细菌性肺炎患者,我们开发了一个自然语言处理(NLP)应用程序来从胸部诊断成像报告中提取相关信息。领域专家通过手动注释282份报告来训练和测试NLP应用程序,从而建立了一个参考标准。胸腔积液、肺浸润物和肺炎的发现被自动从报告中提取出来,然后用于自动分类报告是否与细菌性肺炎一致。与带注释的诊断成像报告参考标准相比,我们的NLP应用程序中最准确的机器学习算法实现允许提取相关结果,灵敏度为.939,阳性预测值为.925。它允许将报告分类,敏感度为0.71,阳性预测值为0.86,特异度为.962。与手动注释这些报告的每个领域专家相比,NLP应用程序允许更高的敏感度(0.71比.527)以及相似的阳性预测值和特异度。在这项初步研究中,基于NLP的肺炎信息提取在儿科诊断成像报告中的表现优于领域专家。NLP是一种从大量影像报告中提取信息的有效方法,有助于CER。
Community-acquired pneumonia is a leading cause of pediatric morbidity. Administrative data are often used to conduct comparative effectiveness research (CER) with sufficient sample sizes to enhance detection of important outcomes. However, such studies are prone to misclassification errors because of the variable accuracy of discharge diagnosis codes. The aim of this study was to develop an automated, scalable, and accurate method to determine the presence or absence of pneumonia in children using chest imaging reports. The multi-institutional PHIS+ clinical repository was developed to support pediatric CER by expanding an administrative database of children’s hospitals with detailed clinical data. To develop a scalable approach to find patients with bacterial pneumonia more accurately, we developed a Natural Language Processing (NLP) application to extract relevant information from chest diagnostic imaging reports. Domain experts established a reference standard by manually annotating 282 reports to train and then test the NLP application. Findings of pleural effusion, pulmonary infiltrate, and pneumonia were automatically extracted from the reports and then used to automatically classify whether a report was consistent with bacterial pneumonia. Compared with the annotated diagnostic imaging reports reference standard, the most accurate implementation of machine learning algorithms in our NLP application allowed extracting relevant findings with a sensitivity of .939 and a positive predictive value of .925. It allowed classifying reports with a sensitivity of .71, a positive predictive value of .86, and a specificity of .962. When compared with each of the domain experts manually annotating these reports, the NLP application allowed for significantly higher sensitivity (.71 vs .527) and similar positive predictive value and specificity . NLP-based pneumonia information extraction of pediatric diagnostic imaging reports performed better than domain experts in this pilot study. NLP is an efficient method to extract information from a large collection of imaging reports to facilitate CER.