Large-scale identification of patients with cerebral aneurysms using natural language processing

Large-scale identification of patients with cerebral aneurysms using natural language processing
复制标题

DOI:
10.1212/wnl.0000000000003490
复制
发表时间:
2017-01-10
期刊:
影响因子:
9.9
通讯作者:
Du, Rose
Du, Rose
中科院分区:
医学1区
文献类型:
--
作者:
Castro, Victor M.;Dligach, Dmitriy;Du, Rose

文献摘要

被引文献

相似文献

目的:使用自然语言处理(NLP)结合电子病历(EMR)来准确识别脑动脉瘤患者及其匹配对照。方法:使用 ICD-9 和当前程序术语代码从 EMR 中获取潜在动脉瘤患者的初始数据集市。然后使用 NLP 来训练分类算法,并使用 0.632 引导交叉验证来校正过度拟合偏差。然后将分类规则应用于完整的数据集市。对 300 名患有动脉瘤的患者进行了额外验证。通过匹配年龄、性别、种族和医疗保健用途来获得对照。结果:我们在 420 万名患者中确定了 55,675 名患者,其 ICD-9 和当前程序术语代码与脑动脉瘤一致。其中,16,823 名患者的术语“动脉瘤”出现在相关解剖学术语附近。训练后,选择了由 8 个编码变量和 14 个 NLP 变量组成的最终算法,得出接收者操作特征曲线下的总面积为 0.95。应用最终算法后,5,589 名患者被分类为患有动脉瘤,54,952 名对照者与这些患者进行了匹配。基于 300 名患者的验证队列的阳性预测值为 0.86。结论:我们通过应用 NLP 来利用 EMR 的力量来获得大量颅内动脉瘤患者及其匹配的对照。这种算法可以推广到其他疾病的流行病学和遗传学研究。神经病学(R)2017; 88:164-168
Objective: To use natural language processing (NLP) in conjunction with the electronic medical record (EMR) to accurately identify patients with cerebral aneurysms and their matched controls.Methods: ICD-9 and Current Procedural Terminology codes were used to obtain an initial data mart of potential aneurysm patients from the EMR. NLP was then used to train a classification algorithm with .632 bootstrap cross-validation used for correction of overfitting bias. The classification rule was then applied to the full data mart. Additional validation was performed on 300 patients classified as having aneurysms. Controls were obtained by matching age, sex, race, and healthcare use.Results: We identified 55,675 patients of 4.2 million patients with ICD-9 and Current Procedural Terminology codes consistent with cerebral aneurysms. Of those, 16,823 patients had the term aneurysm occur near relevant anatomic terms. After training, a final algorithm consisting of 8 coded and 14 NLP variables was selected, yielding an overall area under the receiver-operating characteristic curve of 0.95. After the final algorithm was applied, 5,589 patients were classified as having aneurysms, and 54,952 controls were matched to those patients. The positive predictive value based on a validation cohort of 300 patients was 0.86.Conclusions: We harnessed the power of the EMR by applying NLP to obtain a large cohort of patients with intracranial aneurysms and their matched controls. Such algorithms can be generalized to other diseases for epidemiologic and genetic studies. Neurology (R) 2017; 88:164-168