Top scoring pairs for feature selection in machine learning and applications to cancer outcome prediction.

Top scoring pairs for feature selection in machine learning and applications to cancer outcome prediction.
复制标题

DOI:
10.1186/1471-2105-12-375
复制
发表时间:
2011-09-23
期刊:
影响因子:
3
通讯作者:
Kon MA
Kon MA
中科院分区:
生物学4区
文献类型:
--
作者:
Shi P;Ray S;Zhu Q;Kon MA

文献摘要

参考文献

被引文献

相似文献

广泛使用的k最高评分对(k-TSP)算法是一种简单而强大的无参数分类器。它在许多癌症微阵列数据集上的成功归功于一种有效的基于基因对相对表达排序的特征选择算法。然而,它的一般鲁棒性不扩展到一些困难的数据集,例如涉及癌症结果预测的数据集,这可能是由于分类器使用的相对简单的投票方案。我们认为,可以通过分离其有效的特征选择组件,并将其与一个强大的分类器,如支持向量机(SVM)相结合,以提高性能。更一般地,k-TSP排名算法生成的最高评分对可以用作其他机器学习分类器的降维子空间。我们开发了一种将k-TSP排名算法(TSP)与其他机器学习方法集成的方法,允许将k-TSP的计算效率,多变量特征排名与多变量分类器(如SVM)相结合。我们评估了这种混合方案(K-TSP+SVM)在一系列的模拟数据集与已知的数据结构。与其他特征选择方法,如一个单变量的方法类似于Fisher的判别准则(Fisher),或递归特征消除嵌入SVM(RFE)相比,TSP是越来越有效的比其他两种方法的信息基因变得越来越相关,这是证明在分类性能和恢复真正的信息基因的能力。我们还将这种混合方案应用于四个癌症预后数据集,其中k-TSP+SVM在所有数据集中的性能优于k-TSP分类器,并实现与单独使用SVM相当或上级的性能。与模拟中观察到的一致,TSP在一些癌症数据集中似乎是比Fisher和RFE更好的特征选择器。k-TSP排名算法可以用作机器学习中特征选择的计算效率高的多变量过滤方法。在模拟数据集和一些癌症预后数据集中,SVM与k-TSP排序算法的组合优于k-TSP和SVM单独使用。模拟研究表明,作为一个功能选择器,它是更好地调整到某些数据特征,即信息基因之间的相关性,这是潜在的有趣的替代功能排序方法在途径分析。
The widely used k top scoring pair (k-TSP) algorithm is a simple yet powerful parameter-free classifier. It owes its success in many cancer microarray datasets to an effective feature selection algorithm that is based on relative expression ordering of gene pairs. However, its general robustness does not extend to some difficult datasets, such as those involving cancer outcome prediction, which may be due to the relatively simple voting scheme used by the classifier. We believe that the performance can be enhanced by separating its effective feature selection component and combining it with a powerful classifier such as the support vector machine (SVM). More generally the top scoring pairs generated by the k-TSP ranking algorithm can be used as a dimensionally reduced subspace for other machine learning classifiers. We developed an approach integrating the k-TSP ranking algorithm (TSP) with other machine learning methods, allowing combination of the computationally efficient, multivariate feature ranking of k-TSP with multivariate classifiers such as SVM. We evaluated this hybrid scheme (k-TSP+SVM) in a range of simulated datasets with known data structures. As compared with other feature selection methods, such as a univariate method similar to Fisher's discriminant criterion (Fisher), or a recursive feature elimination embedded in SVM (RFE), TSP is increasingly more effective than the other two methods as the informative genes become progressively more correlated, which is demonstrated both in terms of the classification performance and the ability to recover true informative genes. We also applied this hybrid scheme to four cancer prognosis datasets, in which k-TSP+SVM outperforms k-TSP classifier in all datasets, and achieves either comparable or superior performance to that using SVM alone. In concurrence with what is observed in simulation, TSP appears to be a better feature selector than Fisher and RFE in some of the cancer datasets The k-TSP ranking algorithm can be used as a computationally efficient, multivariate filter method for feature selection in machine learning. SVM in combination with k-TSP ranking algorithm outperforms k-TSP and SVM alone in simulated datasets and in some cancer prognosis datasets. Simulation studies suggest that as a feature selector, it is better tuned to certain data characteristics, i.e. correlations among informative genes, which is potentially interesting as an alternative feature ranking method in pathway analysis.
DOI: 10.1016/s0140-6736(05)17947-1
发表时间: 2005-02-19
期刊: LANCET
影响因子: 168.9
作者:
Wang, YX;Klijn, JGM;Foekens, JA
通讯作者: Foekens, JA
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1093/bioinformatics/bti033
发表时间: 2005-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Statnikov, A;Aliferis, CF;Levy, S
通讯作者: Levy, S
DOI: 10.1073/pnas.0903931106
发表时间: 2009-06-02
影响因子: 11.1
作者:
Jin, Jiashun
通讯作者: Jin, Jiashun
DOI: 10.1093/bioinformatics/bti192
发表时间: 2005-04-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Wang, YH;Makedon, FS;Pearlman, J
通讯作者: Pearlman, J