Granular SVM with Repetitive Undersampling for Highly Imbalanced Protein Homology Prediction

Granular SVM with Repetitive Undersampling for Highly Imbalanced Protein Homology Prediction
复制标题

DOI:
10.1109/grc.2006.1635839
复制
发表时间:
2006-05
期刊:
2006 IEEE International Conference on Granular Computing
影响因子:
--
通讯作者:
Yuchun Tang;Yanqing Zhang
Yuchun Tang;Yanqing Zhang
中科院分区:
其他
文献类型:
--
作者:
Yuchun Tang;Yanqing Zhang

文献摘要

被引文献

相似文献

随着包括生物医学信息学在内的新机器学习应用领域的出现,高度不平衡的分类非常重要并且越来越普遍。为了解决这个具有挑战性的类不平衡问题,本文设计了一种新颖的粒度支持向量机 - 重复欠采样算法(GSVM-RU)。 GSVM-RU创造性地利用支持向量机(SVM)本身进行欠采样,以最大限度地减少信息丢失的负面影响,同时最大限度地提高欠采样过程中数据清理的积极作用。因此,可以对准确且快速的分类器进行建模。 GSVM-RU 在 ACM KDDCUP 2004 竞赛中因极度不平衡的蛋白质同源性预测而被评为最佳解决方案之一。
Highly imbalanced classification is important and increasingly common with emergence of new machine learning application domains including biomedical informatics. In order to solve this challenging class imbalance problem, a novel Granular Support Vector Machines - Repetitive Undersampling algorithm (GSVM-RU) is designed in this work. GSVM-RU creatively utilizes Support Vector Machines (SVM) themselves for undersampling to minimize the negative effect of information loss while maximizing the positive effect of data cleaning in the undersampling process. Consequently, an accurate and fast classifier can be modeled. GSVM-RU ranks as one of the best solutions in ACM KDDCUP 2004 competition for the extremely imbalanced protein homology prediction.