Variance Component Estimation for Mixed Model Analysis of cDNA Microarray Data

Variance Component Estimation for Mixed Model Analysis of cDNA Microarray Data
复制标题

DOI:
10.1002/bimj.200810476
复制
发表时间:
2008-12-01
影响因子:
1.7
通讯作者:
Piepho, Hans-Peter
Piepho, Hans-Peter
中科院分区:
生物学3区
文献类型:
--
作者:
Sarholz, Barbara;Piepho, Hans-Peter

文献摘要

被引文献

相似文献

微阵列为基因表达的定量提供了有价值的工具。然而,通常,在基因混合模型分析中,重复次数有限导致方差估计值不令人满意。由于可获得数千个基因,因此期望跨基因联合收割机组合信息。当要比较两种以上的组织类型或治疗时,建议将阵列效应视为随机效应。然后,可以恢复阵列之间的信息,这可以提高估计的准确性。提出了一种双效应线性混合模型跨基因方差分量估计的方法。该方法可以扩展到具有两个以上随机效应的模型。我们假设方差分量服从对数正态分布。假设来自基因分析的平方和,给定真实的方差分量,遵循比例卡方(2)分布,我们采用经验贝叶斯方法。方差分量由其后验分布的期望估计。新方法进行了评估,在模拟研究。基于这些方差估计的检验比基于基因方差估计的检验更有可能检测到差异表达的基因。这种效应在具有小at射线数的研究中最为明显。对玉米胚乳真实的数据集的分析表明,该方法效果良好。
Microarrays provide a valuable tool for the quantification of gene expression. Usually, however, there is a limited number of replicates leading to unsatisfying variance estimates in a gene-wise mixed model analysis. As thousands of genes are available, it is desirable to combine information across genes. When more than two tissue types or treatments are to be compared it might be advisable to consider the array effect as random. Then information between arrays may be recovered, which can increase accuracy in estimation. We propose a method of variance component estimation across genes for a linear mixed model with two random effects. The method may be extended to models with more than two random effects. We assume that the variance components follow a log-normal distribution. Assuming that the sums of squares from the gene-wise analysis, given the true variance components, follow a scaled chi(2)-distribution, we adopt an empirical Bayes approach. The variance components are estimated by the expectation of their posterior distribution. The new method is evaluated in a simulation study. Differentially expressed genes are more likely to be detected by tests based on these variance estimates than by tests based on gene-wise variance estimates. This effect is most visible in studies with small at-ray numbers. Analyzing a real data set on maize endosperm the method is shown to work well.