Optimal combination of feature selection and classification via local hyperplane based learning strategy.

Optimal combination of feature selection and classification via local hyperplane based learning strategy.
复制标题

通过基于局部超平面的学习策略来优化特征选择和分类的组合

DOI:
10.1186/s12859-015-0629-6
复制
发表时间:
2015-07-10
期刊:
影响因子:
3
通讯作者:
Su W
Su W
中科院分区:
生物学4区
文献类型:
--
作者:
Cheng X;Cai H;Zhang Y;Xu B;Su W

文献摘要

参考文献

被引文献

相似文献

通过基因选择对癌症进行分类是生物医学中最重要且最具挑战性的步骤之一。一个主要的挑战是设计一种有效的方法,从分类中消除不相关的,冗余的,或嘈杂的基因,同时保留所有的高度区分的基因。我们提出了一种基因选择方法,称为局部超平面判别分析(LHDA)。LHDA采用了两个中心思想。首先,它使用局部近似而不是全局测量;其次,它将最近报道的分类模型,K-局部超平面距离最近邻(HKNN)分类器嵌入到其网络中。通过基于分类精度的迭代,LHDA得到特征权重向量,最终提取最优特征子集。在合成和真实的微阵列基准数据集上进行了广泛的实验,评估了所提出的方法的性能。本文采用了八种经典的特征选择方法、四种分类模型和两种流行的嵌入式学习方法,包括k-最近邻(KNN)、超平面k-最近邻(HKNN)、支持向量机(SVM)和随机森林。所提出的方法产生了相当或上级性能七个国家的最先进的模型。良好的性能证明了将特征加权和模型学习结合到一个统一的框架中以同时实现这两个任务的优越性。本文的在线版本(doi:10.1186/s12859-015-0629-6)包含补充材料,可供授权用户使用。
Classifying cancers by gene selection is among the most important and challenging procedures in biomedicine. A major challenge is to design an effective method that eliminates irrelevant, redundant, or noisy genes from the classification, while retaining all of the highly discriminative genes. We propose a gene selection method, called local hyperplane-based discriminant analysis (LHDA). LHDA adopts two central ideas. First, it uses a local approximation rather than global measurement; second, it embeds a recently reported classification model, K-Local Hyperplane Distance Nearest Neighbor(HKNN) classifier, into its discriminator. Through classification accuracy-based iterations, LHDA obtains the feature weight vector and finally extracts the optimal feature subset. The performance of the proposed method is evaluated in extensive experiments on synthetic and real microarray benchmark datasets. Eight classical feature selection methods, four classification models and two popular embedded learning schemes, including k-nearest neighbor (KNN), hyperplane k-nearest neighbor (HKNN), Support Vector Machine (SVM) and Random Forest are employed for comparisons. The proposed method yielded comparable to or superior performances to seven state-of-the-art models. The nice performance demonstrate the superiority of combining feature weighting with model learning into an unified framework to achieve the two tasks simultaneously. The online version of this article (doi:10.1186/s12859-015-0629-6) contains supplementary material, which is available to authorized users.
DOI: 10.1186/1471-2105-7-228
发表时间: 2006-04-27
期刊: BMC bioinformatics
影响因子: 3
作者:
Yang K;Cai Z;Li J;Lin G
通讯作者: Lin G
DOI: 10.1002/wics.101
发表时间: 2010-07-01
影响因子: 3.2
作者:
Abdi, Herve;Williams, Lynne J.
通讯作者: Williams, Lynne J.
DOI: 10.1109/tpami.2005.55
发表时间: 2005-03-01
影响因子: 23.6
作者:
He, XF;Yan, SC;Zhang, HJ
通讯作者: Zhang, HJ
DOI: 10.1109/tpami.2007.250598
发表时间: 2007-01-01
影响因子: 23.6
作者:
Yan, Shuicheng;Xu, Dong;Lin, Stephen
通讯作者: Lin, Stephen
DOI: 10.6026/97320630004385
发表时间: 2010-02-28
期刊: Bioinformation
影响因子: 1.9
作者:
Pok G;Liu JC;Ryu KH
通讯作者: Ryu KH