Structuring clinical text with AI: Old versus new natural language processing techniques evaluated on eight common cardiovascular diseases.

Structuring clinical text with AI: Old versus new natural language processing techniques evaluated on eight common cardiovascular diseases.
复制标题

DOI:
10.1016/j.patter.2021.100289
复制
发表时间:
2021-07-09
期刊:
Patterns (New York, N.Y.)
影响因子:
--
通讯作者:
Gevaert O
Gevaert O
中科院分区:
其他
文献类型:
--
作者:
Zhan X;Humbert-Droz M;Mukherjee P;Gevaert O

文献摘要

参考文献

被引文献

相似文献

电子健康记录中的自由文本临床笔记对于数据挖掘来说更困难,而结构化诊断代码可能缺失或有误。为了提高诊断代码的质量,这项工作从自由文本笔记中提取诊断代码:使用了五种新旧词向量化方法对斯坦福病程记录进行向量化,并通过逻辑回归预测八种常见心血管疾病的ICD - 10代码。这些模型表现良好,其中TF - IDF作为最佳向量化模型,显示出最高的AUROC(0.9499 - 0.9915)和AUPRC(0.2956 - 0.8072)。当在MIMIC - III数据上进行测试时,这些模型也显示出可迁移性,AUROC为0.7952到0.9790,AUPRC为0.2353到0.8084。通过与每种疾病匹配的具有临床意义的重要词汇展示了模型的可解释性。这项研究表明了从自由文本临床笔记中准确提取结构化诊断代码、填补缺失代码以及纠正错误代码以用于信息检索和下游机器学习应用的可行性。 五种自然语言处理词向量化模型以高AUROC和AUPRC预测8种ICD - 10代码 表现最佳的TF - IDF模型通过重要词汇显示出完全的可解释性 当在MIMIC - III重症监护病房数据集上进行测试时,这些模型显示出高可迁移性 电子健康记录中结构化数据(如诊断代码)的挖掘使许多临床应用成为可能,但大量临床信息被锁定在非结构化的自由文本临床笔记中,因为它们在数据挖掘中更难使用。此外,结构化诊断代码经常缺失甚至有误。为了将自由文本笔记准确地以诊断代码的形式结构化以便下游使用,我们使用新旧自然语言处理方法以及可解释的分类算法来提取八种常见心血管疾病的诊断代码。这项工作有助于对自由文本临床笔记进行结构化,填补缺失的诊断代码,并纠正临床医生记录的错误诊断代码,以提高诊断代码作为后续信息检索和下游数据挖掘应用的基础结构化数据的数据质量。 电子健康记录中的非结构化数据对于数据挖掘来说很困难,而结构化诊断代码经常缺失甚至有误。这项工作使用五种新旧自然语言处理技术来提取八种心血管疾病的诊断代码,这有助于对自由文本临床笔记进行结构化,填补缺失的诊断代码,并纠正临床医生记录的错误诊断代码,以提高诊断代码作为后续信息检索和下游数据挖掘应用的基础结构化数据的数据质量。
Free-text clinical notes in electronic health records are more difficult for data mining while the structured diagnostic codes can be missing or erroneous. To improve the quality of diagnostic codes, this work extracts diagnostic codes from free-text notes: five old and new word vectorization methods were used to vectorize Stanford progress notes and predict eight ICD-10 codes of common cardiovascular diseases with logistic regression. The models showed good performance, with TF-IDF as the best vectorization model showing the highest AUROC (0.9499–0.9915) and AUPRC (0.2956–0.8072). The models also showed transferability when tested on MIMIC-III data with AUROC from 0.7952 to 0.9790 and AUPRC from 0.2353 to 0.8084. Model interpretability was shown by the important words with clinical meanings matching each disease. This study shows the feasibility of accurately extracting structured diagnostic codes, imputing missing codes, and correcting erroneous codes from free-text clinical notes for information retrieval and downstream machine-learning applications. Five NLP word vectorization models predict 8 ICD-10 codes with high AUROC and AUPRC The best-performing TF-IDF models showed full interpretability with important words The models showed high transferability when tested on the MIMIC-III ICU dataset The mining of the structured data in electronic health records such as diagnostic codes enables many clinical applications, but much clinical information is locked in the unstructured free-text clinical notes because they are more difficult to use in data mining. In addition, the structured diagnostic codes are often missing or even erroneous. To accurately structure the free-text notes in the form of diagnostic code for downstream usage, we used old and new natural language processing methods together with interpretable classification algorithms to extract eight diagnostic codes of common cardiovascular diseases. This work helps to structure free-text clinical notes, impute missing diagnostic codes, and correct erroneously diagnostic codes noted by clinicians to improve the data quality of diagnostic codes as the fundamental structured data for later information retrieval and downstream data-mining applications. Unstructured data in electronic health records are difficult for data mining while structured diagnostic codes are often missing or even erroneous. This work used five old and new natural language processing techniques to extract eight cardiovascular diseases' diagnostic codes, which helps to structure free-text clinical notes, impute missing diagnostic codes, and correct erroneously diagnostic codes noted by clinicians to improve the data quality of diagnostic codes as the fundamental structured data for later information retrieval and downstream data-mining applications.
DOI: 10.1016/j.cmpb.2019.05.024
发表时间: 2019-08-01
影响因子: 6.1
作者:
Huang, Jinmiao;Osorio, Cesar;Sy, Luke Wicent
通讯作者: Sy, Luke Wicent
DOI: 10.1016/j.ebiom.2019.06.034
发表时间: 2019-07-01
期刊: EBIOMEDICINE
影响因子: 11.1
作者:
Huang, Chao;Cintra, Murilo;Gevaert, Olivier
通讯作者: Gevaert, Olivier
DOI: 10.1161/jaha.115.003056
发表时间: 2016-05-31
影响因子: 5.4
作者:
Chang TE;Lichtman JH;Goldstein LB;George MG
通讯作者: George MG
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG
DOI: 10.1186/s13023-018-0830-6
发表时间: 2018-05-31
影响因子: 3.7
作者:
Garcelon N;Neuraz A;Salomon R;Bahi-Buisson N;Amiel J;Picard C;Mahlaoui N;Benoit V;Burgun A;Rance B
通讯作者: Rance B