A Review of Matched-pairs Feature Selection Methods for Gene Expression Data Analysis.

A Review of Matched-pairs Feature Selection Methods for Gene Expression Data Analysis.
复制标题

基因表达数据分析的匹配对特征选择方法综述

DOI:
10.1016/j.csbj.2018.02.005
复制
发表时间:
2018
影响因子:
6
通讯作者:
Ma Q
Ma Q
中科院分区:
生物学2区
文献类型:
--
作者:
Liang S;Ma A;Yang S;Wang Y;Ma Q

文献摘要

参考文献

被引文献

相似文献

随着微阵列、rna测序(RNA-seq)、单细胞RNA-seq等技术对基因表达数据的快速积累,有必要对这些高维数据进行降维和特征(签名基因)选择,以支持对这些高维数据的理解。这些计算方法极大地促进了进一步的数据分析和解释,如基因功能富集分析、癌症生物标志物检测和精准医学中的药物靶向鉴定。尽管生物信息学中已经开发了许多特征选择方法,但如何针对特定问题选择合适的方法并寻求最合理的排序特征仍然是一个挑战。与此同时,配对病例对照设计(matching case-control design, MCCD)下的配对基因表达数据越来越受欢迎,常用于多组学整合研究,可以通过抵消混淆特征的相似分布来提高特征选择效率。然而,针对配对数据设计的合适的特征选择方法,即匹配-配对特征选择(MPFS),并没有得到成熟的并行发展。本文采用3种分类方法,比较了10种特征选择方法(8种MPFS方法和2种传统非配对方法)在2个真实数据集上的性能,并通过程序运行分析了这些方法的算法复杂度。本文旨在归纳和全面地呈现MPFS,以便读者容易理解其特征,并在选择合适的分析方法时获得线索。
With the rapid accumulation of gene expression data from various technologies, e.g., microarray, RNA-sequencing (RNA-seq), and single-cell RNA-seq, it is necessary to carry out dimensional reduction and feature (signature genes) selection in support of making sense out of such high dimensional data. These computational methods significantly facilitate further data analysis and interpretation, such as gene function enrichment analysis, cancer biomarker detection, and drug targeting identification in precision medicine. Although numerous methods have been developed for feature selection in bioinformatics, it is still a challenge to choose the appropriate methods for a specific problem and seek for the most reasonable ranking features. Meanwhile, the paired gene expression data under matched case-control design (MCCD) is becoming increasingly popular, which has often been used in multi-omics integration studies and may increase feature selection efficiency by offsetting similar distributions of confounding features. The appropriate feature selection methods specifically designed for the paired data, which is named as matched-pairs feature selection (MPFS), however, have not been maturely developed in parallel. In this review, we compare the performance of 10 feature-selection methods (eight MPFS methods and two traditional unpaired methods) on two real datasets by applied three classification methods, and analyze the algorithm complexity of these methods through the running of their programs. This review aims to induce and comprehensively present the MPFS in such a way that readers can easily understand its characteristics and get a clue in selecting the appropriate methods for their analyses.
DOI: 10.1214/09-ejs537
发表时间: 2009-01-01
影响因子: 1.1
作者:
Buneo, Florentino;Barbu, Adrian
通讯作者: Barbu, Adrian
DOI: 10.1038/srep10312
发表时间: 2015-05-19
期刊: Scientific reports
影响因子: 4.6
作者:
Bermingham ML;Pong-Wong R;Spiliopoulou A;Hayward C;Rudan I;Campbell H;Wright AF;Wilson JF;Agakov F;Navarro P;Haley CS
通讯作者: Haley CS
DOI: 10.1198/016214503000125
发表时间: 2003-06-01
影响因子: 3.7
作者:
Bühlmann, P;Yu, B
通讯作者: Yu, B
视觉识别中的稀疏表示和学习:理论与应用
DOI: 10.1016/j.sigpro.2012.09.011
发表时间: 2013-06-01
期刊: SIGNAL PROCESSING
影响因子: 4.4
作者:
Cheng, Hong;Liu, Zicheng;Chen, Xuewen
通讯作者: Chen, Xuewen
DOI: 10.1007/978-1-60327-194-3_11
发表时间: 2010-01-01
期刊: BIOINFORMATICS METHODS IN CLINICAL RESEARCH
影响因子: --
作者:
Datta, Susmita;Pihur, Vasyl
通讯作者: Pihur, Vasyl