Natural language processing algorithms for mapping clinical text fragments onto ontology concepts: a systematic review and recommendations for future studies.

Natural language processing algorithms for mapping clinical text fragments onto ontology concepts: a systematic review and recommendations for future studies.
复制标题

将临床文本片段映射到本体概念的自然语言处理算法:系统性综述和对未来研究的建议。

DOI:
10.1186/s13326-020-00231-z
复制
发表时间:
2020-11-16
影响因子:
1.9
通讯作者:
Arts DL
Arts DL
中科院分区:
工程技术4区
文献类型:
--
作者:
Kersloot MG;van Putten FJP;Abu-Hanna A;Cornet R;Arts DL

文献摘要

参考文献

被引文献

相似文献

电子健康记录(EHR)中的自由文本描述可能对临床研究和护理优化感兴趣。然而,自由文本不能容易地被计算机解释,因此,具有有限的价值。自然语言处理(NLP)算法可以通过在自由文本中附加本体概念使其变得可机器解释,然而,NLP算法的实现并没有得到一致的评价。因此,本研究的目的是回顾目前用于开发和评估NLP算法的方法,将临床文本片段映射到本体概念。为了标准化算法的评估并减少研究之间的异质性,我们提出了一系列建议。两名评审员检查了Scopus、IEEE、MEDLINE、EMBASE、ACM数字图书馆和ACL Anthology索引的出版物。其中包括报告NLP将临床文本从EHR映射到本体概念的出版物。提取了年份、国家、设置、目标、评估和验证方法、NLP算法、术语系统、数据集大小和语言、性能指标、参考标准、可推广性、操作使用和源代码可用性。研究目的以归纳法分类。这些结果被用来确定建议。确定了2355项独特的研究。256项研究报告了将自由文本映射到本体概念的NLP算法的发展。77份报告介绍了发展和评价情况。22项研究未对未知数据进行验证,68项研究未进行外部验证。在23项声称他们的算法是可推广的研究中,有5项通过外部验证进行了测试。关于NLP系统和算法的使用,数据的使用,评估和验证,结果的呈现和结果的普遍性的十六个建议的列表被开发。我们发现了许多异构的方法来报告NLP算法的开发和评估,这些算法将临床文本映射到本体概念。超过四分之一的已识别出版物未进行评价。此外,超过四分之一的纳入研究没有进行验证,88%没有进行外部验证。我们相信,我们的建议以及现有的报告标准将提高未来研究和医学中NLP算法的可重复性和可重用性。补充信息随附于10.1186/s13326-020-00231-z。
Free-text descriptions in electronic health records (EHRs) can be of interest for clinical research and care optimization. However, free text cannot be readily interpreted by a computer and, therefore, has limited value. Natural Language Processing (NLP) algorithms can make free text machine-interpretable by attaching ontology concepts to it. However, implementations of NLP algorithms are not evaluated consistently. Therefore, the objective of this study was to review the current methods used for developing and evaluating NLP algorithms that map clinical text fragments onto ontology concepts. To standardize the evaluation of algorithms and reduce heterogeneity between studies, we propose a list of recommendations. Two reviewers examined publications indexed by Scopus, IEEE, MEDLINE, EMBASE, the ACM Digital Library, and the ACL Anthology. Publications reporting on NLP for mapping clinical text from EHRs to ontology concepts were included. Year, country, setting, objective, evaluation and validation methods, NLP algorithms, terminology systems, dataset size and language, performance measures, reference standard, generalizability, operational use, and source code availability were extracted. The studies’ objectives were categorized by way of induction. These results were used to define recommendations. Two thousand three hundred fifty five unique studies were identified. Two hundred fifty six studies reported on the development of NLP algorithms for mapping free text to ontology concepts. Seventy-seven described development and evaluation. Twenty-two studies did not perform a validation on unseen data and 68 studies did not perform external validation. Of 23 studies that claimed that their algorithm was generalizable, 5 tested this by external validation. A list of sixteen recommendations regarding the usage of NLP systems and algorithms, usage of data, evaluation and validation, presentation of results, and generalizability of results was developed. We found many heterogeneous approaches to the reporting on the development and evaluation of NLP algorithms that map clinical text to ontology concepts. Over one-fourth of the identified publications did not perform an evaluation. In addition, over one-fourth of the included studies did not perform a validation, and 88% did not perform external validation. We believe that our recommendations, alongside an existing reporting standard, will increase the reproducibility and reusability of future studies and NLP algorithms in medicine. Supplementary information accompanies this paper at 10.1186/s13326-020-00231-z.
DOI: 10.1136/amiajnl-2014-002954
发表时间: 2015-04-01
影响因子: 6.4
作者:
Bejan, Cosmin Adrian;Wei, Wei-Qi;Denny, Joshua C.
通讯作者: Denny, Joshua C.
DOI: 10.1016/j.ijmedinf.2019.04.022
发表时间: 2019-07-01
影响因子: 4.9
作者:
Becker, Matthias;Kasper, Stefan;Virchow, Isabel
通讯作者: Virchow, Isabel
Stard 2015:用于报告诊断准确性研究的基本项目的更新列表。
DOI: 10.1136/bmj.h5527
发表时间: 2015-10-28
期刊: BMJ (Clinical research ed.)
影响因子: --
作者:
Bossuyt PM;Reitsma JB;Bruns DE;Gatsonis CA;Glasziou PP;Irwig L;Lijmer JG;Moher D;Rennie D;de Vet HC;Kressel HY;Rifai N;Golub RM;Altman DG;Hooft L;Korevaar DA;Cohen JF;STARD Group
通讯作者: STARD Group
DOI: 10.1016/j.aca.2012.11.007
发表时间: 2013-01-14
影响因子: 6.2
作者:
Beleites, Claudia;Neugebauer, Ute;Popp, Juergen
通讯作者: Popp, Juergen