Nearest Neighbor-Based Strategy to Optimize Multi-View Triplet Network for Classification of Small-Sample Medical Imaging Data

Nearest Neighbor-Based Strategy to Optimize Multi-View Triplet Network for Classification of Small-Sample Medical Imaging Data
复制标题

DOI:
10.1109/tnnls.2021.3059635
复制
发表时间:
2021-03
影响因子:
10.4
通讯作者:
Phawis Thammasorn;W. Chaovalitwongse;D. Hippe;L. Wootton;Eric Ford;M. Spraker;S. Combs;J. Peeken;Matthew Nyflot
Phawis Thammasorn;W. Chaovalitwongse;D. Hippe;L. Wootton;Eric Ford;M. Spraker;S. Combs;J. Peeken;Matthew Nyflot
中科院分区:
计算机科学1区
文献类型:
--
作者:
Phawis Thammasorn;W. Chaovalitwongse;D. Hippe;L. Wootton;Eric Ford;M. Spraker;S. Combs;J. Peeken;Matthew Nyflot

文献摘要

相似文献

有限样本和数据扩充的多视图分类是医学中非常常见的机器学习(ML)问题。在有限数据的情况下,提出了一种用于两阶段表征学习的三元网络方法。然而,有效的训练和验证的代表性网络的功能,他们在后续的分类器的适用性仍然是未解决的问题。尽管用于训练的典型的基于距离的度量捕获了特征的整体类可分性,但是根据这些度量的性能并不总是导致最佳分类。因此,需要对所有特征-分类器组合进行详尽的调优以搜索最佳最终结果。为了克服这一挑战,我们开发了一种新的最近邻(NN)验证策略的基础上的三重度量。该策略由理论基础支持,以提供具有最高端性能的下限的特征的最佳选择。所提出的策略是一个透明的方法来确定是否要改善的功能或分类。这避免了重复调谐的需要。我们对真实世界医学成像任务的评估(即,放射治疗递送误差预测和肉瘤存活预测)表明我们的策略上级其他常见的深度表示学习基线[即,自动编码器(AE)和softmax]。该策略解决了特征的可解释性问题,从而实现更全面的特征创建,使得医学专家可以专注于指定相关数据,而不是繁琐的特征工程。
Multi-view classification with limited sample size and data augmentation is a very common machine learning (ML) problem in medicine. With limited data, a triplet network approach for two-stage representation learning has been proposed. However, effective training and verifying the features from the representation network for their suitability in subsequent classifiers are still unsolved problems. Although typical distance-based metrics for the training capture the overall class separability of the features, the performance according to these metrics does not always lead to an optimal classification. Consequently, an exhaustive tuning with all feature–classifier combinations is required to search for the best end result. To overcome this challenge, we developed a novel nearest-neighbor (NN) validation strategy based on the triplet metric. This strategy is supported by a theoretical foundation to provide the best selection of the features with a lower bound of the highest end performance. The proposed strategy is a transparent approach to identify whether to improve the features or the classifier. This avoids the need for repeated tuning. Our evaluations on real-world medical imaging tasks (i.e., radiation therapy delivery error prediction and sarcoma survival prediction) show that our strategy is superior to other common deep representation learning baselines [i.e., autoencoder (AE) and softmax]. The strategy addresses the issue of feature’s interpretability which enables more holistic feature creation such that the medical experts can focus on specifying relevant data as opposed to tedious feature engineering.