Detecting differential gene expression with a semiparametric hierarchical mixture method

Detecting differential gene expression with a semiparametric hierarchical mixture method
复制标题

DOI:
10.1093/biostatistics/5.2.155
复制
发表时间:
2004-04-01
期刊:
影响因子:
2.1
通讯作者:
Ahlquist, P
Ahlquist, P
中科院分区:
数学2区
文献类型:
--
作者:
Newton, MA;Noueiry, A;Ahlquist, P

文献摘要

被引文献

相似文献

混合模型为微阵列数据分析中的差异表达问题提供了一种有效的方法。基于完全参数混合模型的方法是可用的,但在某些示例中缺乏拟合表明更灵活的模型可能是有益的。现有的更灵活的混合模型在一维基因特异性汇总统计量的水平上工作,因此当每个基因的测量相对较少时,这些方法可能无法提供差异表达的灵敏检测器。我们提出了一个分层混合模型,提供的方法,是敏感的检测差异表达和足够灵活的解释归一化微阵列数据的复杂变异性。基于EM的算法被用来拟合模型的参数和半参数版本。我们将注意力限制在两个样本的比较问题上,一个涉及Affyphine微阵列和酵母翻译的实验提供了一个激励性的案例研究。差异表达的基因特异性后验概率形成了统计推断的基础;它们定义了短基因列表和错误发现率。与几种竞争的方法相比,所提出的方法在模拟研究中表现出良好的操作特性,对扣球数据的分析,并在交叉验证计算。
Mixture modeling provides an effective approach to the differential expression problem in microarray data analysis. Methods based on fully parametric mixture models are available, but lack of fit in some examples indicates that more flexible models may be beneficial. Existing, more flexible, mixture models work at the level of one-dimensional gene-specific summary statistics, and so when there are relatively few measurements per gene these methods may not provide sensitive detectors of differential expression. We propose a hierarchical mixture model to provide methodology that is both sensitive in detecting differential expression and sufficiently flexible to account for the complex variability of normalized microarray data. EM-based algorithms are used to fit both parametric and semiparametric versions of the model. We restrict attention to the two-sample comparison problem; an experiment involving Affymetrix microarrays and yeast translation provides the motivating case study. Gene-specific posterior probabilities of differential expression form the basis of statistical inference; they define short gene lists and false discovery rates. Compared to several competing methodologies, the proposed methodology exhibits good operating characteristics in a simulation study, on the analysis of spike-in data, and in a cross-validation calculation.