Can Sequential Images from the Same Object Be Used for Training Machine Learning Models? A Case Study for Detecting Liver Disease by Ultrasound Radiomics.

Can Sequential Images from the Same Object Be Used for Training Machine Learning Models? A Case Study for Detecting Liver Disease by Ultrasound Radiomics.
复制标题

DOI:
10.3390/ai3030043
复制
发表时间:
2022-09
期刊:
AI (Basel, Switzerland)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

被引文献

相似文献

用于医学成像的机器学习不仅需要足够的数据来进行训练和测试,而且数据是独立的。当观测值之间存在内在相关性时,经常会看到高度相互依赖的数据。这对于从时间序列中获取的连续成像数据尤其是可以预期的。在这项研究中,我们评估了使用统计措施来测试从同一情况下采取的连续超声图像数据的独立性。共分析了1180例B型肝脏超声图像,共5903个感兴趣区域。超声图像取自两个肝脏疾病组,纤维化和脂肪变性,以及正常病例。计算机提取的纹理特征,然后用来训练机器学习(ML)模型的计算机辅助诊断。该实验导致使用逻辑回归的高两类诊断,AUC为0.928,以及使用随机森林ML的高多类分类性能,AUC为0.917。为了评估机器学习的图像区域独立性,使用了Jenson-Shannon(JS)发散。JS分布表明,正常肝脏的图像相互独立,而两种疾病病理的图像并不独立。为了保证机器学习模型的通用性,并防止数据泄漏,在机器学习之前,应测试获取的同一对象的多帧图像数据的独立性。这样的测试可以应用于现实世界的医学图像问题,以确定来自同一对象的图像是否可以用于训练。
Machine learning for medical imaging not only requires sufficient amounts of data for training and testing but also that the data be independent. It is common to see highly interdependent data whenever there are inherent correlations between observations. This is especially to be expected for sequential imaging data taken from time series. In this study, we evaluate the use of statistical measures to test the independence of sequential ultrasound image data taken from the same case. A total of 1180 B-mode liver ultrasound images with 5903 regions of interests were analyzed. The ultrasound images were taken from two liver disease groups, fibrosis and steatosis, as well as normal cases. Computer-extracted texture features were then used to train a machine learning (ML) model for computer-aided diagnosis. The experiment resulted in high two-category diagnosis using logistic regression, with AUC of 0.928 and high performance of multicategory classification, using random forest ML, with AUC of 0.917. To evaluate the image region independence for machine learning, Jenson–Shannon (JS) divergence was used. JS distributions showed that images of normal liver were independent from each other, while the images from the two disease pathologies were not independent. To guarantee the generalizability of machine learning models, and to prevent data leakage, multiple frames of image data acquired of the same object should be tested for independence before machine learning. Such tests can be applied to real-world medical image problems to determine if images from the same subject can be used for training.