Modeling considerations for using expression data from multiple species.

Modeling considerations for using expression data from multiple species.
复制标题

使用来自多个物种的表达数据的建模注意事项。

DOI:
10.1002/sim.5850
复制
发表时间:
2013
影响因子:
2
通讯作者:
Kechris,KaterinaJ
Kechris,KaterinaJ
中科院分区:
医学3区
文献类型:
--
作者:
Siewert,Elizabeth;Kechris,KaterinaJ

文献摘要

相似文献

尽管来自多个物种的全基因组表达数据集现在更普遍地生成,但关于如何最好地将这种类型的相关数据整合到模型中的研究很少。从预测转录因子结合位点的单物种线性回归模型作为案例研究开始,我们研究了如何在将该模型扩展到多个物种时最好地考虑相关表达数据。使用多元回归模型,我们以两种方式解释了物种之间的系统发育关系:(i)重复测量模型,其中误差项受到约束;(ii)贝叶斯层次模型,其中回归系数的先验分布受到约束。我们表明,两种多物种模型都比单物种模型提高了预测性能。当相互比较时,重复测量模型优于贝叶斯模型。我们提出了一个可能的解释与约束误差项的模型更好的性能。版权所有©2013 John Wiley & Sons, Ltd
Although genome‐wide expression data sets from multiple species are now more commonly generated, there have been few studies on how to best integrate this type of correlated data into models. Starting with a single‐species, linear regression model that predicts transcription factor binding sites as a case study, we investigated how best to take into account the correlated expression data when extending this model to multiple species. Using a multivariate regression model, we accounted for the phylogenetic relationships among the species in two ways: (i) a repeated‐measures model, where the error term is constrained; and (ii) a Bayesian hierarchical model, where the prior distributions of the regression coefficients are constrained. We show that both multiple‐species models improve predictive performance over the single‐species model. When compared with each other, the repeated‐measures model outperformed the Bayesian model. We suggest a possible explanation for the better performance of the model with the constrained error term. Copyright © 2013 John Wiley & Sons, Ltd.