A CNN-based unified framework utilizing projection loss in unison with label noise handling for multiple Myeloma cancer diagnosis

A CNN-based unified framework utilizing projection loss in unison with label noise handling for multiple Myeloma cancer diagnosis
复制标题

DOI:
10.1016/j.media.2021.102099
复制
发表时间:
2021-06-04
影响因子:
10.9
通讯作者:
Gupta, Ritu
Gupta, Ritu
中科院分区:
工程技术1区
文献类型:
--
作者:
Gehlot, Shiv;Gupta, Anubha;Gupta, Ritu

文献摘要

被引文献

相似文献

多发性骨髓瘤(MM)是浆细胞恶性肿瘤。与其他形式的癌症类似,它需要及时诊断以降低死亡风险。传统的诊断工具是资源密集型的,因此,这些解决方案不容易扩展以将其范围扩展到大众。深度学习的进步导致了经济实惠、资源优化、易于部署的计算机辅助解决方案的快速发展。这项工作提出了一个统一的框架,MM诊断使用显微镜血细胞成像数据,解决健康与癌细胞的类间视觉相似性和数据集的标签噪声的关键挑战。为了提取类的区别性特征,我们提出了投影损失,以最大限度地提高投影的样本的激活各自的类向量,除了施加正交约束的类向量。该投影损失与交叉熵损失一起沿着用于设计双分支架构,其有助于实现改进的性能并提供针对标签噪声问题的范围。基于这种架构,已经提出了两种方法来校正噪声标签。一个耦合分类器也被提出来解决双分支架构的预测中的冲突。我们利用了72名受试者(26名健康受试者和46名MM癌症受试者)的大型数据集,共包含74996张图像(包括34555张训练细胞图像和40441张测试细胞图像)。这是迄今为止文献中报道的关于多发性骨髓瘤癌症的最广泛的数据集。还进行了消融研究。建议的架构表现最好的平衡精度为94。17%的健康与癌症的二进制细胞分类与十个最先进的架构的比较性能。两个额外的公开可用的数据集的两种不同的方式也被用于分析所提出的方法的标签噪声处理能力的广泛的实验。该代码将在https://github.com/shivgahlout/CAD-MM上提供。(c)2021爱思唯尔有限公司版权所有。
Multiple Myeloma (MM) is a malignancy of plasma cells. Similar to other forms of cancer, it demands prompt diagnosis for reducing the risk of mortality. The conventional diagnostic tools are resource-intense and hence, these solutions are not easily scalable for extending their reach to the masses. Advancements in deep learning have led to rapid developments in affordable, resource optimized, easily deployable computer-assisted solutions. This work proposes a unified framework for MM diagnosis using microscopic blood cell imaging data that addresses the key challenges of inter-class visual similarity of healthy versus cancer cells and that of the label noise of the dataset. To extract class distinctive features, we propose projection loss to maximize the projection of a sample's activation on the respective class vector besides imposing orthogonality constraints on the class vectors. This projection loss is used along with the cross-entropy loss to design a dual branch architecture that helps achieve improved performance and provides scope for targeting the label noise problem. Based on this architecture, two methodologies have been proposed to correct the noisy labels. A coupling classifier has also been proposed to resolve the conflicts in the dual-branch architecture's predictions. We have utilized a large dataset of 72 subjects (26 healthy and 46 MM cancer) containing a total of 74996 images (including 34555 training cell images and 40441 test cell images). This is so far the most extensive dataset on Multiple Myeloma cancer ever reported in the literature. An ablation study has also been carried out. The proposed architecture performs best with a balanced accuracy of 94 . 17% on binary cell classification of healthy versus cancer in the comparative performance with ten state-of-the-art architectures. Extensive experiments on two additional publicly available datasets of two different modalities have also been utilized for analyzing the label noise handling capability of the proposed methodology. The code will be available under https://github.com/shivgahlout/CAD-MM . (c) 2021 Elsevier B.V. All rights reserved.