Dissecting High-Dimensional Phenotypes with Bayesian Sparse Factor Analysis of Genetic Covariance Matrices

Dissecting High-Dimensional Phenotypes with Bayesian Sparse Factor Analysis of Genetic Covariance Matrices
复制标题

DOI:
10.1534/genetics.113.151217
复制
发表时间:
2013-07-01
期刊:
影响因子:
3.3
通讯作者:
Mukherjee, Sayan
Mukherjee, Sayan
中科院分区:
生物学2区
文献类型:
--
作者:
Runcie, Daniel E.;Mukherjee, Sayan

文献摘要

被引文献

相似文献

模拟复杂、多变量表型的定量遗传研究对于进化预测和人工选择都很重要。例如,基因表达的变化可以提供对连接基因型和表型的发育和生理机制的洞察。然而,经典的分析技术不太适合基因表达的定量遗传学研究,其中每个个体测定的性状数量可以达到数千个。在这里,我们推导出一个贝叶斯遗传稀疏因子模型估计的遗传协方差矩阵(G矩阵)的高维性状,如基因表达,在混合效应模型。我们的模型的关键思想是,我们只需要考虑生物学上合理的G-矩阵。生物体的整个表型是模块化过程的结果,具有有限的复杂性。这意味着G矩阵将是高度结构化的。特别是,我们假设中间特征(或因素,例如,发育或生理学上的变异)控制着高维表型的变异,并且这些中间性状中的每一个都是稀疏的-仅影响少数观察到的性状。这种方法的优点有两方面。首先,稀疏因子是可解释的,并提供对遗传结构潜在机制的生物学见解。其次,强制稀疏性有助于防止采样错误淹没高维数据中的真实信号。我们证明了我们的模型的优势,在模拟数据和已发表的果蝇基因表达数据集的分析。
Quantitative genetic studies that model complex, multivariate phenotypes are important for both evolutionary prediction and artificial selection. For example, changes in gene expression can provide insight into developmental and physiological mechanisms that link genotype and phenotype. However, classical analytical techniques are poorly suited to quantitative genetic studies of gene expression where the number of traits assayed per individual can reach many thousand. Here, we derive a Bayesian genetic sparse factor model for estimating the genetic covariance matrix (G-matrix) of high-dimensional traits, such as gene expression, in a mixed-effects model. The key idea of our model is that we need consider only G-matrices that are biologically plausible. An organism's entire phenotype is the result of processes that are modular and have limited complexity. This implies that the G-matrix will be highly structured. In particular, we assume that a limited number of intermediate traits (or factors, e.g., variations in development or physiology) control the variation in the high-dimensional phenotype, and that each of these intermediate traits is sparse - affecting only a few observed traits. The advantages of this approach are twofold. First, sparse factors are interpretable and provide biological insight into mechanisms underlying the genetic architecture. Second, enforcing sparsity helps prevent sampling errors from swamping out the true signal in high-dimensional data. We demonstrate the advantages of our model on simulated data and in an analysis of a published Drosophila melanogaster gene expression data set.