T-staging pulmonary oncology from radiological reports using natural language processing: translating into a multi-language setting.

T-staging pulmonary oncology from radiological reports using natural language processing: translating into a multi-language setting.
复制标题

DOI:
10.1186/s13244-021-01018-1
复制
发表时间:
2021-06-10
影响因子:
4.7
通讯作者:
Dekker ALAJ
Dekker ALAJ
中科院分区:
医学2区
文献类型:
--
作者:
Nobel JM;Puts S;Weiss J;Aerts HJWL;Mak RH;Robben SGF;Dekker ALAJ

文献摘要

参考文献

被引文献

相似文献

在数字化时代,医疗数据的准确性和结构化对于多种应用非常重要。特别是肿瘤分期的数据需要准确,以分期和治疗患者,以及人群水平的监测和结果评估。为了支持从自由文本放射学报告中提取数据,建立了荷兰自然语言处理(NLP)算法,以根据肿瘤淋巴结转移(TNM)分类来量化肺肿瘤的T分期。该结构化工具已在英语放射学自由文本报告中进行了翻译和验证。一个基于规则的算法来分类T-阶段进行了训练和验证,分别为200和225英语自由文本放射学报告从诊断计算机断层扫描(CT)分期肺癌患者。将通过算法从报告中提取的自动T分期与手动分期进行比较。一个图形用户界面的训练目的,可视化的算法的结果,突出显示提取的概念及其修改上下文。T分期分类器的准确度在验证集中为0.89,考虑T亚期时为0.84,仅考虑肿瘤大小时为0.76。结果与荷兰的结果相当(分别为0.88、0.89和0.79)。大多数错误是由于模糊性问题造成的,这些问题无法通过基于规则的算法来解决。NLP可以成功地应用于不同语言的自由文本放射学报告中的肺癌分期。应采用混合方法重点引入机器学习,以提高性能。在线版本包含补充材料,可通过10.1186/s13244-021-01018-1获得。
In the era of datafication, it is important that medical data are accurate and structured for multiple applications. Especially data for oncological staging need to be accurate to stage and treat a patient, as well as population-level surveillance and outcome assessment. To support data extraction from free-text radiological reports, Dutch natural language processing (NLP) algorithm was built to quantify T-stage of pulmonary tumors according to the tumor node metastasis (TNM) classification. This structuring tool was translated and validated on English radiological free-text reports. A rule-based algorithm to classify T-stage was trained and validated on, respectively, 200 and 225 English free-text radiological reports from diagnostic computed tomography (CT) obtained for staging of patients with lung cancer. The automated T-stage extracted by the algorithm from the report was compared to manual staging. A graphical user interface was built for training purposes to visualize the results of the algorithm by highlighting the extracted concepts and its modifying context. Accuracy of the T-stage classifier was 0.89 in the validation set, 0.84 when considering the T-substages, and 0.76 when only considering tumor size. Results were comparable with the Dutch results (respectively, 0.88, 0.89 and 0.79). Most errors were made due to ambiguity issues that could not be solved by the rule-based nature of the algorithm. NLP can be successfully applied for staging lung cancer from free-text radiological reports in different languages. Focused introduction of machine learning should be introduced in a hybrid approach to improve performance. The online version contains supplementary material available at 10.1186/s13244-021-01018-1.
DOI: 10.1186/s13244-020-00907-1
发表时间: 2020-09-29
影响因子: 4.7
作者:
Weber TF;Spurny M;Hasse FC;Sedlaczek O;Haag GM;Springfeld C;Mokry T;Jäger D;Kauczor HU;Berger AK
通讯作者: Berger AK
DOI: 10.1001/jama.243.8.756
发表时间: 1980-01-01
影响因子: 120.7
作者:
COTE, RA;ROBBOY, S
通讯作者: ROBBOY, S
DOI: 10.2214/ajr.14.12636
发表时间: 2014-12-01
影响因子: 5
作者:
Marcovici, Peter A.;Taylor, George A.
通讯作者: Taylor, George A.
DOI: 10.1016/j.jbi.2011.03.011
发表时间: 2011-10
影响因子: 4.5
作者:
Chapman BE;Lee S;Kang HP;Chapman WW
通讯作者: Chapman WW
DOI: 10.1007/s13244-013-0218-z
发表时间: 2013-04-01
影响因子: 4.7
作者:
通讯作者: --