Robust gene selection methods using weighting schemes for microarray data analysis.

Robust gene selection methods using weighting schemes for microarray data analysis.
复制标题

DOI:
10.1186/s12859-017-1810-x
复制
发表时间:
2017-09-02
期刊:
影响因子:
3
通讯作者:
Song J
Song J
中科院分区:
生物学4区
文献类型:
--
作者:
Kang S;Song J

文献摘要

参考文献

被引文献

相似文献

微阵列数据分析中的一个常见任务是识别在两种不同状态之间差异表达的信息基因。由于微阵列数据的高维性质,识别重要基因在分析数据中是必不可少的。然而,许多基因选择技术的性能高度依赖于实验条件,例如存在测量误差或有限数量的样品重复。我们提出了新的基于过滤器的基因选择技术,通过应用一个简单的修改,显着性分析的微阵列(SAM)。为了证明所提出的方法的有效性,我们考虑了一系列的合成数据集与不同的噪声水平和样本大小沿着与两个真实的数据集。调查结果如下。首先,我们提出的方法优于传统方法的所有模拟设置。特别是,当给定的数据是噪声和样本量小,我们的方法要好得多。他们表现出相对稳健的性能,无论噪声水平和样本量,而SAM的性能变得显着变差的噪声水平变得高或样本量减少。当有足够的样品重复时,SAM和我们的方法表现出相似的性能。最后,我们提出的方法是有竞争力的传统方法在分类任务的微阵列。仿真研究和真实的数据分析的结果表明,本文提出的方法对于检测显著基因和分类任务是有效的,特别是当给定的数据有噪声或样本重复数很少时。通过采用加权方案,我们可以获得强大的和可靠的结果,微阵列数据分析。本文的在线版本(10.1186/s12859-017-1810-x)包含补充材料,可供授权用户使用。
A common task in microarray data analysis is to identify informative genes that are differentially expressed between two different states. Owing to the high-dimensional nature of microarray data, identification of significant genes has been essential in analyzing the data. However, the performances of many gene selection techniques are highly dependent on the experimental conditions, such as the presence of measurement error or a limited number of sample replicates. We have proposed new filter-based gene selection techniques, by applying a simple modification to significance analysis of microarrays (SAM). To prove the effectiveness of the proposed method, we considered a series of synthetic datasets with different noise levels and sample sizes along with two real datasets. The following findings were made. First, our proposed methods outperform conventional methods for all simulation set-ups. In particular, our methods are much better when the given data are noisy and sample size is small. They showed relatively robust performance regardless of noise level and sample size, whereas the performance of SAM became significantly worse as the noise level became high or sample size decreased. When sufficient sample replicates were available, SAM and our methods showed similar performance. Finally, our proposed methods are competitive with traditional methods in classification tasks for microarrays. The results of simulation study and real data analysis have demonstrated that our proposed methods are effective for detecting significant genes and classification tasks, especially when the given data are noisy or have few sample replicates. By employing weighting schemes, we can obtain robust and reliable results for microarray data analysis. The online version of this article (10.1186/s12859-017-1810-x) contains supplementary material, which is available to authorized users.
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1002/sim.2109
发表时间: 2005-08-15
影响因子: 2
作者:
Kooperberg, C;Aragaki, A;Olson, JM
通讯作者: Olson, JM
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1002/cfg.62
发表时间: 2001
影响因子: --
作者:
Dougherty, E R
通讯作者: Dougherty, E R
DOI: 10.1007/s13748-015-0080-y
发表时间: 2016-05-01
影响因子: 4.2
作者:
Bolon-Canedo, Veronica;Sanchez-Marono, Noelia;Alonso-Betanzos, Amparo
通讯作者: Alonso-Betanzos, Amparo