Natural Language Processing of Radiology Reports in Patients With Hepatocellular Carcinoma to Predict Radiology Resource Utilization

Natural Language Processing of Radiology Reports in Patients With Hepatocellular Carcinoma to Predict Radiology Resource Utilization
复制标题

DOI:
10.1016/j.jacr.2018.12.004
复制
发表时间:
2019-06-01
影响因子:
4.5
通讯作者:
Kachura, J. R.
Kachura, J. R.
中科院分区:
医学3区
文献类型:
--
作者:
Brown, A. D.;Kachura, J. R.

文献摘要

被引文献

相似文献

目的:放射学是一种有限的卫生保健资源,在大多数卫生中心都有很高的需求。然而,由于疾病预测的内在不确定性,预测需求波动是一项挑战。本研究的目的是探讨自然语言处理(NLP)在预测肝细胞癌监测患者下游放射资源利用方面的潜力。材料和方法:2010年1月1日至2017年10月31日在我所进行的所有肝细胞癌监测CT检查均选自我科放射学信息系统。我们使用开源的自然语言处理和机器学习软件将放射学报告文本解析为词袋和术语频率倒置文档频率(TF-IDF)表示。三种机器学习模型--Logistic回归、支持向量机和随机森林被用来预测放射科资源的未来利用。测试数据集被用来计算准确度、敏感度和特异度以及曲线下面积(AUC)。结果:词袋模型总体上略逊于TF-IDF特征提取方法。TF-IDF+支持向量机模型的准确率为92%,敏感度为83%,特异度为96%,AUC为0.971。结论:基于NLP的模型可以从叙述性的肝癌监测报告中准确地预测下游放射学资源的利用,并有可能转化为医疗保健管理,从而改善决策,降低成本,扩大获得医疗保健的机会。
Objective: Radiology is a finite health care resource in high demand at most health centers. However, anticipating fluctuations in demand is a challenge because of the inherent uncertainty in disease prognosis. The aim of this study was to explore the potential of natural language processing (NLP) to predict downstream radiology resource utilization in patients undergoing surveillance for hepatocellular carcinoma (HCC).Materials and Methods: All HCC surveillance CT examinations performed at our institution from January 1, 2010, to October 31, 2017 were selected from our departmental radiology information system. We used open source NLP and machine learning software to parse radiology report text into bag-of-words and term frequency-inverse document frequency (TF-IDF) representations. Three machine learning models-logistic regression, support vector machine (SVM), and random forest-were used to predict future utilization of radiology department resources. A test data set was used to calculate accuracy, sensitivity, and specificity in addition to the area under the curve (AUC).Results: As a group, the bag-of-word models were slightly inferior to the TF-IDF feature extraction approach. The TF-IDF + SVM model outperformed all other models with an accuracy of 92%, a sensitivity of 83%, and a specificity of 96%, with an AUC of 0.971.Conclusions: NLP-based models can accurately predict downstream radiology resource utilization from narrative HCC surveillance reports and has potential for translation to health care management where it may improve decision making, reduce costs, and broaden access to care.