Classification of imbalanced oral cancer image data from high-risk population.

Classification of imbalanced oral cancer image data from high-risk population.
复制标题

DOI:
10.1117/1.jbo.26.10.105001
复制
发表时间:
2021-10
影响因子:
3.5
通讯作者:
Liang R
Liang R
中科院分区:
医学3区
文献类型:
--
作者:
Song B;Li S;Sunny S;Gurushanth K;Mendonca P;Mukhia N;Patrick S;Gurudath S;Raghavan S;Tsusennaro I;Leivon ST;Kolur T;Shetty V;Bushan V;Ramesh R;Peterson T;Pillai V;Wilder-Smith P;Sigamani A;Suresh A;Kuriakose MA;Birur P;Liang R

文献摘要

被引文献

相似文献

重要性:口腔癌的早期检测对于高危患者至关重要,基于机器学习的自动分类是疾病筛查的理想选择。然而,目前从高风险人群中收集的数据集是不平衡的,往往对分类的性能产生不利影响。目的:减少数据不平衡引起的类偏差。方法:我们收集了3851偏振白色光颊粘膜图像使用我们定制的口腔癌筛查设备。我们使用权重平衡,数据增强,欠采样,焦点损失和集成方法来改善口腔癌图像分类的神经网络性能,这些图像分类使用了在低资源环境下口腔癌筛查期间从高危人群中捕获的不平衡多类数据集。结果如下:通过将数据级和算法级方法应用于深度学习训练过程,一开始难以区分的少数类的性能得到了改善。“癌前病变”类别的准确性也得到提高,这是筛查应用的理想选择。结论:实验结果表明,不平衡的口腔癌图像数据集引起的类偏差可以减少使用数据和算法级的方法。我们的研究可能为帮助理解不平衡数据集对口腔癌深度学习分类器的影响以及如何减轻提供了重要基础。
Significance: Early detection of oral cancer is vital for high-risk patients, and machine learning-based automatic classification is ideal for disease screening. However, current datasets collected from high-risk populations are unbalanced and often have detrimental effects on the performance of classification. Aim: To reduce the class bias caused by data imbalance. Approach: We collected 3851 polarized white light cheek mucosa images using our customized oral cancer screening device. We use weight balancing, data augmentation, undersampling, focal loss, and ensemble methods to improve the neural network performance of oral cancer image classification with the imbalanced multi-class datasets captured from high-risk populations during oral cancer screening in low-resource settings. Results: By applying both data-level and algorithm-level approaches to the deep learning training process, the performance of the minority classes, which were difficult to distinguish at the beginning, has been improved. The accuracy of “premalignancy” class is also increased, which is ideal for screening applications. Conclusions: Experimental results show that the class bias induced by imbalanced oral cancer image datasets could be reduced using both data- and algorithm-level methods. Our study may provide an important basis for helping understand the influence of unbalanced datasets on oral cancer deep learning classifiers and how to mitigate.