LICIC: Less Important Components for Imbalanced Multiclass Classification

LICIC: Less Important Components for Imbalanced Multiclass Classification
复制标题

DOI:
10.3390/info9120317
复制
发表时间:
2018-12-01
期刊:
影响因子:
3.1
通讯作者:
Pirlo, Giuseppe
Pirlo, Giuseppe
中科院分区:
其他
文献类型:
--
作者:
Dentamaro, Vincenzo;Impedovo, Donato;Pirlo, Giuseppe

文献摘要

被引文献

相似文献

癌症诊断中的多类分类,使用DNA或基因表达签名,以及MALDI-TOF质谱数据中的细菌物种指纹的分类,由于不平衡的数据和相对于实例数量的高维度数量而具有挑战性。在这项研究中,一个新的过采样技术称为LICIC将作为一个有价值的工具,在对付两个类的不平衡,和著名的灾难的维数问题。该方法能够保留数据集中的非线性,同时创建新实例而不增加噪音。该方法将与其他过采样方法进行比较,如随机过采样,SMOTE,Borderline-SMOTE和ADASYN。F1分数显示了这种新技术在使用不平衡,多类和高维数据集时的有效性。
Multiclass classification in cancer diagnostics, using DNA or Gene Expression Signatures, but also classification of bacteria species fingerprints in MALDI-TOF mass spectrometry data, is challenging because of imbalanced data and the high number of dimensions with respect to the number of instances. In this study, a new oversampling technique called LICIC will be presented as a valuable instrument in countering both class imbalance, and the famous curse of dimensionality problem. The method enables preservation of non-linearities within the dataset, while creating new instances without adding noise. The method will be compared with other oversampling methods, such as Random Oversampling, SMOTE, Borderline-SMOTE, and ADASYN. F1 scores show the validity of this new technique when used with imbalanced, multiclass, and high-dimensional datasets.