Early recognition of multiple sclerosis using natural language processing of the electronic health record.

Early recognition of multiple sclerosis using natural language processing of the electronic health record.
复制标题

使用电子健康记录的自然语言处理来早期识别多发性硬化症。

DOI:
10.1186/s12911-017-0418-4
复制
发表时间:
2017-02-28
影响因子:
3.5
通讯作者:
Fulgieri DJ
Fulgieri DJ
中科院分区:
医学3区
文献类型:
--
作者:
Chase HS;Mitrani LR;Lu GG;Fulgieri DJ

文献摘要

被引文献

相似文献

通过在电子健康记录(EHR)中搜索患者的临床笔记以寻找多发性硬化症(MS)等疾病的体征和症状,诊断的准确性可能会得到提高。这项研究的重点是确定在医疗保健提供者最初识别之前,是否可以从他们的临床笔记中识别出MS患者。来自成人门诊的MS患者(n = 165例)和对照组(n = 545例)均为MS富集组。随机样本队列来自同一成人门诊随机选择的患者(n = 2289),其中一些人患有MS(n = 16)。患者的笔记从数据仓库中提取,并使用MedLEE将体征和症状映射到UMLS术语。大约1000个与多发性硬化症相关的词汇在多发性硬化症患者笔记中的出现频率显著高于对照组。同义词被手动聚类到50个桶中,用作分类特征。采用朴素贝叶斯分类将患者分为MS组和非MS组。使用初始ICD9[MS]编码后输入的MS丰富队列的笔记对已知的MS患者进行分类,得到的ROC AUC、灵敏度和特异度分别为0.90[0.87-0.93]、0.75[0.66-0.82]和0.91[0.87-0.93]。使用随机样本队列中的音符也获得了类似的分类精度。使用在最初的ICD9[MS]文档之前输入的富含MS的队列的笔记对未知患有MS的患者进行分类发现40%[23-59%]患有MS手动回顾随机样本队列中被归类为患有MS但缺乏ICD9[MS]代码的45名患者的EHR确定4名可能不认识MS的患者通过使用NLP挖掘患者的临床笔记以了解特定疾病的体征和症状,可能会提高诊断的准确性。使用这种方法,我们在疾病病程的早期识别出患有多发性硬化症的患者,这可能会缩短诊断时间。这种方法也可以应用于其他经常被初级保健提供者遗漏的疾病,如癌症。实施计算机化的诊断支持最终是否会缩短从最早的症状到正式识别这种疾病的时间,还有待观察。
Diagnostic accuracy might be improved by algorithms that searched patients’ clinical notes in the electronic health record (EHR) for signs and symptoms of diseases such as multiple sclerosis (MS). The focus this study was to determine if patients with MS could be identified from their clinical notes prior to the initial recognition by their healthcare providers. An MS-enriched cohort of patients with well-established MS (n = 165) and controls (n = 545), was generated from the adult outpatient clinic. A random sample cohort was generated from randomly selected patients (n = 2289) from the same adult outpatient clinic, some of whom had MS (n = 16). Patients’ notes were extracted from the data warehouse and signs and symptoms mapped to UMLS terms using MedLEE. Approximately 1000 MS-related terms occurred significantly more frequently in MS patients’ notes than controls’. Synonymous terms were manually clustered into 50 buckets and used as classification features. Patients were classified as MS or not using Naïve Bayes classification. Classification of patients known to have MS using notes of the MS-enriched cohort entered after the initial ICD9[MS] code yielded an ROC AUC, sensitivity, and specificity of 0.90 [0.87-0.93], 0.75[0.66-0.82], and 0.91 [0.87-0.93], respectively. Similar classification accuracy was achieved using the notes from the random sample cohort. Classification of patients not yet known to have MS using notes of the MS-enriched cohort entered before the initial ICD9[MS] documentation identified 40% [23–59%] as having MS. Manual review of the EHR of 45 patients of the random sample cohort classified as having MS but lacking an ICD9[MS] code identified four who might have unrecognized MS. Diagnostic accuracy might be improved by mining patients’ clinical notes for signs and symptoms of specific diseases using NLP. Using this approach, we identified patients with MS early in the course of their disease which could potentially shorten the time to diagnosis. This approach could also be applied to other diseases often missed by primary care providers such as cancer. Whether implementing computerized diagnostic support ultimately shortens the time from earliest symptoms to formal recognition of the disease remains to be seen.