Structuring clinical text with AI: Old versus new natural language processing techniques evaluated on eight common cardiovascular diseases.
Structuring clinical text with AI: Old versus new natural language processing techniques evaluated on eight common cardiovascular diseases.
复制标题
DOI:
10.1016/j.patter.2021.100289
复制
发表时间:
2021-07-09
期刊:
影响因子:
--
通讯作者:
Gevaert O
中科院分区:
文献类型:
--
作者:
Zhan X;Humbert-Droz M;Mukherjee P;Gevaert O
Free-text clinical notes in electronic health records are more difficult for data mining while the structured diagnostic codes can be missing or erroneous. To improve the quality of diagnostic codes, this work extracts diagnostic codes from free-text notes: five old and new word vectorization methods were used to vectorize Stanford progress notes and predict eight ICD-10 codes of common cardiovascular diseases with logistic regression. The models showed good performance, with TF-IDF as the best vectorization model showing the highest AUROC (0.9499–0.9915) and AUPRC (0.2956–0.8072). The models also showed transferability when tested on MIMIC-III data with AUROC from 0.7952 to 0.9790 and AUPRC from 0.2353 to 0.8084. Model interpretability was shown by the important words with clinical meanings matching each disease. This study shows the feasibility of accurately extracting structured diagnostic codes, imputing missing codes, and correcting erroneous codes from free-text clinical notes for information retrieval and downstream machine-learning applications. Five NLP word vectorization models predict 8 ICD-10 codes with high AUROC and AUPRC The best-performing TF-IDF models showed full interpretability with important words The models showed high transferability when tested on the MIMIC-III ICU dataset The mining of the structured data in electronic health records such as diagnostic codes enables many clinical applications, but much clinical information is locked in the unstructured free-text clinical notes because they are more difficult to use in data mining. In addition, the structured diagnostic codes are often missing or even erroneous. To accurately structure the free-text notes in the form of diagnostic code for downstream usage, we used old and new natural language processing methods together with interpretable classification algorithms to extract eight diagnostic codes of common cardiovascular diseases. This work helps to structure free-text clinical notes, impute missing diagnostic codes, and correct erroneously diagnostic codes noted by clinicians to improve the data quality of diagnostic codes as the fundamental structured data for later information retrieval and downstream data-mining applications. Unstructured data in electronic health records are difficult for data mining while structured diagnostic codes are often missing or even erroneous. This work used five old and new natural language processing techniques to extract eight cardiovascular diseases' diagnostic codes, which helps to structure free-text clinical notes, impute missing diagnostic codes, and correct erroneously diagnostic codes noted by clinicians to improve the data quality of diagnostic codes as the fundamental structured data for later information retrieval and downstream data-mining applications.
登录
查看更多内容
影响因子:
6.1
作者:
Huang, Jinmiao;Osorio, Cesar;Sy, Luke Wicent
通讯作者:
Sy, Luke Wicent
影响因子:
11.1
作者:
Huang, Chao;Cintra, Murilo;Gevaert, Olivier
通讯作者:
Gevaert, Olivier
影响因子:
5.4
作者:
Chang TE;Lichtman JH;Goldstein LB;George MG
通讯作者:
George MG
影响因子:
9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者:
Mark RG
影响因子:
3.7
作者:
Garcelon N;Neuraz A;Salomon R;Bahi-Buisson N;Amiel J;Picard C;Mahlaoui N;Benoit V;Burgun A;Rance B
通讯作者:
Rance B