Computational strategies for alternative single-step Bayesian regression models with large numbers of genotyped and non-genotyped animals

Computational strategies for alternative single-step Bayesian regression models with large numbers of genotyped and non-genotyped animals
复制标题

DOI:
10.1186/s12711-016-0273-2
复制
发表时间:
2016-12-08
影响因子:
4.1
通讯作者:
Garrick, Dorian J.
Garrick, Dorian J.
中科院分区:
生物学2区
文献类型:
--
作者:
Fernando, Rohan L.;Cheng, Hao;Garrick, Dorian J.

文献摘要

被引文献

相似文献

背景资料:两种类型的模型已用于单步基因组预测和全基因组关联研究,包括来自基因分型动物及其非基因分型亲属的表型。这两种类型是明确拟合育种值的育种值模型(BVM)和根据观察或估算的基因型的影响表达育种值的标记效应模型(MEM)。MEM可以适应更广泛的分析类别,包括变量选择或混合模型分析。需要解决的方程的顺序和在其建设中所需的逆变化很大,因此所需的计算工作量取决于系谱的大小,基因型动物的数量和locus.Theory的数量:我们提出的计算策略,以避免存储大,密集块的MME,涉及插补基因型。此外,我们提出了一个混合模型,适合MEM的动物观察到的基因型和BVM的那些没有基因型。混合模型是计算上有吸引力的系谱文件包含数百万的动物与大比例的那些被genotyped.Application:我们证明了原始MEM和混合模型的实用性,使用真实的数据与6,179,960动物的系谱与4,934,101表型和31,453动物基因分型在40,214个信息位点。为了完成一个单一性状的分析在台式电脑上有四个图形卡需要约3小时,使用混合模型,以获得预处理共轭梯度的解决方案和42,000马尔可夫链蒙特-卡罗(MCMC)样本的育种值,这使得推断后验均值,方差和协方差。MCMC采样需要四分之一的努力时,使用的混合动力模型相比,公布MEM.Conclusions:我们提出了一个混合动力模型,适合MEM的动物基因型和BVM的那些没有基因型。它的实用性和计算工作量的显着减少被证明。该模型可以很容易地扩展到适应多性状,多品种,母体效应,以及额外的随机效应,如多基因残留效应。
Background: Two types of models have been used for single-step genomic prediction and genome-wide association studies that include phenotypes from both genotyped animals and their non-genotyped relatives. The two types are breeding value models (BVM) that fit breeding values explicitly and marker effects models (MEM) that express the breeding values in terms of the effects of observed or imputed genotypes. MEM can accommodate a wider class of analyses, including variable selection or mixture model analyses. The order of the equations that need to be solved and the inverses required in their construction vary widely, and thus the computational effort required depends upon the size of the pedigree, the number of genotyped animals and the number of loci.Theory: We present computational strategies to avoid storing large, dense blocks of the MME that involve imputed genotypes. Furthermore, we present a hybrid model that fits a MEM for animals with observed genotypes and a BVM for those without genotypes. The hybrid model is computationally attractive for pedigree files containing millions of animals with a large proportion of those being genotyped.Application: We demonstrate the practicality on both the original MEM and the hybrid model using real data with 6,179,960 animals in the pedigree with 4,934,101 phenotypes and 31,453 animals genotyped at 40,214 informative loci. To complete a single-trait analysis on a desk-top computer with four graphics cards required about 3 h using the hybrid model to obtain both preconditioned conjugate gradient solutions and 42,000 Markov chain Monte-Carlo (MCMC) samples of breeding values, which allowed making inferences from posterior means, variances and covariances. The MCMC sampling required one quarter of the effort when the hybrid model was used compared to the published MEM.Conclusions: We present a hybrid model that fits a MEM for animals with genotypes and a BVM for those without genotypes. Its practicality and considerable reduction in computing effort was demonstrated. This model can readily be extended to accommodate multiple traits, multiple breeds, maternal effects, and additional random effects such as polygenic residual effects.