Assessing differential expression in two-color microarrays: a resampling-based empirical Bayes approach.

Assessing differential expression in two-color microarrays: a resampling-based empirical Bayes approach.
复制标题

DOI:
10.1371/journal.pone.0080099
复制
发表时间:
2013
期刊:
影响因子:
3.7
通讯作者:
Dye TD
Dye TD
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Li D;Le Pape MA;Parikh NI;Chen WX;Dye TD

文献摘要

参考文献

被引文献

相似文献

微阵列广泛用于检查差异基因表达、鉴定单核苷酸多态性和检测甲基化位点。微阵列数据分析中的多种测试方法旨在控制I型和II型错误率;然而,真实的微阵列数据并不总是符合其分布假设。例如,Smyth 普遍存在的参数方法不能充分适应对正态性假设的违反,导致 I 类错误率过高。微阵列的显着性分析是另一种广泛使用的微阵列数据分析方法,它基于排列检验,对非正态分布数据具有鲁棒性;然而,微阵列显着性分析方法倍数变化标准存在问题,并且由于分析中对照数据集的成分变化,可能会严重改变研究的结论。我们提出了一种新颖的方法,将重采样与经验贝叶斯方法相结合:基于重采样的经验贝叶斯方法。这种方法不仅降低了非正态分布微阵列数据的错误发现率,而且还不受倍数变化阈值的影响,因为不需要选择控制数据集。通过模拟研究,比较了 Smyth 参数方法、微阵列显着性分析和基于重采样的经验贝叶斯方法的灵敏度、特异性、总拒绝率和错误发现率。通过早产甲基化研究说明了每种方法之间的错误发现率控制的差异。结果表明,当数据不呈正态分布时,与 Smyth 参数方法相比,基于重采样的经验贝叶斯方法具有显着更高的特异性和更低的错误发现率。当正态分布和非正态分布数据的显着差异表达基因的比例都较大时,基于重采样的经验贝叶斯方法还提供比微阵列显着性分析方法更高的统计功效。最后,基于重采样的经验贝叶斯方法可推广到下一代测序 RNA-seq 数据分析。
Microarrays are widely used for examining differential gene expression, identifying single nucleotide polymorphisms, and detecting methylation loci. Multiple testing methods in microarray data analysis aim at controlling both Type I and Type II error rates; however, real microarray data do not always fit their distribution assumptions. Smyth's ubiquitous parametric method, for example, inadequately accommodates violations of normality assumptions, resulting in inflated Type I error rates. The Significance Analysis of Microarrays, another widely used microarray data analysis method, is based on a permutation test and is robust to non-normally distributed data; however, the Significance Analysis of Microarrays method fold change criteria are problematic, and can critically alter the conclusion of a study, as a result of compositional changes of the control data set in the analysis. We propose a novel approach, combining resampling with empirical Bayes methods: the Resampling-based empirical Bayes Methods. This approach not only reduces false discovery rates for non-normally distributed microarray data, but it is also impervious to fold change threshold since no control data set selection is needed. Through simulation studies, sensitivities, specificities, total rejections, and false discovery rates are compared across the Smyth's parametric method, the Significance Analysis of Microarrays, and the Resampling-based empirical Bayes Methods. Differences in false discovery rates controls between each approach are illustrated through a preterm delivery methylation study. The results show that the Resampling-based empirical Bayes Methods offer significantly higher specificity and lower false discovery rates compared to Smyth's parametric method when data are not normally distributed. The Resampling-based empirical Bayes Methods also offers higher statistical power than the Significance Analysis of Microarrays method when the proportion of significantly differentially expressed genes is large for both normally and non-normally distributed data. Finally, the Resampling-based empirical Bayes Methods are generalizable to next generation sequencing RNA-seq data analysis.
DOI: 10.1073/pnas.091062498
发表时间: 2001-04-24
影响因子: 11.1
作者:
Tusher, VG;Tibshirani, R;Chu, G
通讯作者: Chu, G
DOI: 10.1214/aos/1176345638
发表时间: 1981-01-01
影响因子: 4.5
作者:
FREEDMAN, DA
通讯作者: FREEDMAN, DA
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1198/016214501753382129
发表时间: 2001-12-01
影响因子: 3.7
作者:
Efron, B;Tibshirani, R;Tusher, V
通讯作者: Tusher, V
DOI: 10.1002/bdra.20770
发表时间: 2011-08
影响因子: --
作者:
Adkins, Ronald M.;Krushkal, Julia;Tylavsky, Frances A.;Thomas, Fridtjof
通讯作者: Thomas, Fridtjof