LVQ-SMOTE - Learning Vector Quantization based Synthetic Minority Over-sampling Technique for biomedical data.

LVQ-SMOTE - Learning Vector Quantization based Synthetic Minority Over-sampling Technique for biomedical data.
复制标题

DOI:
10.1186/1756-0381-6-16
复制
发表时间:
2013-10-02
期刊:
影响因子:
4.5
通讯作者:
Kimura H
Kimura H
中科院分区:
生物学3区
文献类型:
--
作者:
Nakamura M;Kajiwara Y;Otsuka A;Kimura H

文献摘要

参考文献

被引文献

相似文献

针对生物医学数据不平衡分类问题,提出了基于合成少数派过采样技术的过采样方法。然而,现有的过采样方法的结果比最简单的SMOTE略好,有时甚至更差。为了提高SMOTE的有效性,本文提出了一种利用学习向量量化得到的码本进行过采样的新方法。一般来说,即使现有的SMOTE应用于生物医学数据集,其空白特征空间仍然非常大,以至于大多数分类算法在估计类之间的边界时表现不佳。为了解决这个问题,我们的过采样方法生成了比其他SMOTE算法占用更多特征空间的合成样本。简而言之,我们的过采样方法可以通过参考从现实世界数据集中获取的实际样本来生成有用的合成样本。在8个真实不平衡数据集上的实验表明,我们提出的过采样方法在5种标准分类算法中的4种上都优于最简单的SMOTE算法。此外,如果在我们的算法中使用称为MWMOTE的最新SMOTE,则可以看到我们的方法的性能有所提高。在β-turn类型预测数据集上的实验显示了一些在以前的分析中没有看到的重要模式。提出的过采样方法产生有用的合成样本,用于不平衡生物医学数据的分类。此外,本文提出的过采样方法与基本分类算法和现有的过采样方法基本兼容。
Over-sampling methods based on Synthetic Minority Over-sampling Technique (SMOTE) have been proposed for classification problems of imbalanced biomedical data. However, the existing over-sampling methods achieve slightly better or sometimes worse result than the simplest SMOTE. In order to improve the effectiveness of SMOTE, this paper presents a novel over-sampling method using codebooks obtained by the learning vector quantization. In general, even when an existing SMOTE applied to a biomedical dataset, its empty feature space is still so huge that most classification algorithms would not perform well on estimating borderlines between classes. To tackle this problem, our over-sampling method generates synthetic samples which occupy more feature space than the other SMOTE algorithms. Briefly saying, our over-sampling method enables to generate useful synthetic samples by referring to actual samples taken from real-world datasets. Experiments on eight real-world imbalanced datasets demonstrate that our proposed over-sampling method performs better than the simplest SMOTE on four of five standard classification algorithms. Moreover, it is seen that the performance of our method increases if the latest SMOTE called MWMOTE is used in our algorithm. Experiments on datasets for β-turn types prediction show some important patterns that have not been seen in previous analyses. The proposed over-sampling method generates useful synthetic samples for the classification of imbalanced biomedical data. Besides, the proposed over-sampling method is basically compatible with basic classification algorithms and the existing over-sampling methods.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1007/bf00994018
发表时间: 1995-09-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
CORTES, C;VAPNIK, V
通讯作者: VAPNIK, V
DOI: 10.1186/1471-2105-11-167
发表时间: 2010-04-02
期刊: BMC bioinformatics
影响因子: 3
作者:
Yu CY;Chou LC;Chang DT
通讯作者: Chang DT
DOI: 10.1613/jair.953
发表时间: 2002-01-01
影响因子: 5
作者:
Chawla, NV;Bowyer, KW;Kegelmeyer, WP
通讯作者: Kegelmeyer, WP
DOI: 10.1186/1471-2105-11-407
发表时间: 2010-07-31
期刊: BMC bioinformatics
影响因子: 3
作者:
Kountouris P;Hirst JD
通讯作者: Hirst JD