A Filter Feature Selection Method Based on MFA Score and Redundancy Excluding and It's Application to Tumor Gene Expression Data Analysis

A Filter Feature Selection Method Based on MFA Score and Redundancy Excluding and It's Application to Tumor Gene Expression Data Analysis
复制标题

DOI:
10.1007/s12539-015-0272-y
复制
发表时间:
2015-12-01
影响因子:
4.8
通讯作者:
Pang, Zenan
Pang, Zenan
中科院分区:
生物学3区
文献类型:
--
作者:
Li, Jiangeng;Su, Lei;Pang, Zenan

文献摘要

被引文献

相似文献

近年来,特征选择技术在肿瘤基因表达数据分析中得到了广泛应用。提出了一种基于图嵌入的滤波器特征选择方法——边际费雪分析分数(MFA分数),主要是因为它优于费雪分数而得到了广泛的应用。针对基因表达数据的高冗余性,提出了一种新的滤波特征选择技术。它被命名为MFA分数+,是基于MFA分数和冗余排除。我们将其应用于一个人工数据集和8个肿瘤基因表达数据集来选择重要特征,然后使用支持向量机作为分类器对样本进行分类。与MFA评分、t检验和Fisher评分相比,具有更高的分类准确率。
Feature selection techniques have been widely applied to tumor gene expression data analysis in recent years. A filter feature selection method named marginal Fisher analysis score (MFA score) which is based on graph embedding has been proposed, and it has been widely used mainly because it is superior to Fisher score. Considering the heavy redundancy in gene expression data, we proposed a new filter feature selection technique in this paper. It is named MFA score+ and is based on MFA score and redundancy excluding. We applied it to an artificial dataset and eight tumor gene expression datasets to select important features and then used support vector machine as the classifier to classify the samples. Compared with MFA score, t test and Fisher score, it achieved higher classification accuracy.