Implementation of genomic recursions in single-step genomic best linear unbiased predictor for US Holsteins with a large number of genotyped animals

Implementation of genomic recursions in single-step genomic best linear unbiased predictor for US Holsteins with a large number of genotyped animals
复制标题

DOI:
10.3168/jds.2015-10540
复制
发表时间:
2016-03-01
影响因子:
3.5
通讯作者:
Lawlor, T. J.
Lawlor, T. J.
中科院分区:
农林科学1区
文献类型:
--
作者:
Masuda, Y.;Misztal, I.;Lawlor, T. J.

文献摘要

被引文献

相似文献

本研究的目的是开发和评估在单步基因组BLUP中使用递归算法(称为算法for proven and young (APY))计算基因组关系矩阵逆的有效实现。我们在美国荷斯坦的最终得分中验证了超过50万只基因型动物的年轻公牛的基因组预测。表型数据包括7,093,380头美国荷斯坦奶牛的11,626,576个最终分数,以及569,404头动物的基因型。利用所有表型数据计算了2009年没有分类子代的年轻公牛的子代偏差,但2014年至少有30只分类子代。同样的公牛的基因组预测是用单步基因组BLUP计算的,使用到2009年的表型。基于一小部分基因型动物(核心动物)基因组关系矩阵的直接反演,我们计算了基因组关系矩阵的逆(G(APY)(-1)),并扩展了该信息。递归到非核心动物。我们测试了几组核心动物,包括9406头公牛至少有1个分类子代,9406头公牛和1052头分类的公牛,9406头公牛和7422头分类的奶牛,以及5000到30000只动物的随机样本。验证的可靠性是通过对预测的年轻公牛基因组预测的回归子偏差的决定系数来评估的。随机选择的5000只核心动物的信度为0.39,9406头公牛和7422头奶牛作为核心动物的信度为0.45,其余各组的信度为0.44。将2009年的表型截断,并使用预条件共轭梯度求解混合模型方程,公牛定义的核心动物的收敛轮数为1,343;以公牛和母牛来定义,2066;由10000只随机动物定义,最多1629只。在完整的表型数据下,轮次分别减少到858、1299和最多1092轮。对569,404只基因型动物和10,000只核心动物建立G(APY)(-1)需要1.3小时和57gb内存。当核心动物的数量至少为10,000时,APY的验证可靠性达到平台。在核心动物的定义中,APY预测的可靠性差异不大。单步基因组BLUP与APY适用于数以百万计的基因分型动物。
The objectives of this study were to develop and evaluate an efficient implementation in the computation of the inverse of genomic relationship matrix with the recursion algorithm, called the algorithm for proven and young (APY), in single-step genomic BLUP. We validated genomic predictions for young bulls with more than 500,000 genotyped animals in final score for US Holsteins. Phenotypic data included 11,626,576 final scores on 7,093,380 US Holstein cows, and genotypes were available for 569,404 animals. Daughter deviations for young bulls with no classified daughters in 2009, but at least 30 classified daughters in 2014 were computed using all the phenotypic data. Genomic predictions for the same bulls were calculated with singlestep genomic BLUP using phenotypes up to 2009. We calculated the inverse of the genomic relationship matrix (G(APY)(-1)) based on a direct inversion of genomic relationship matrix on a small subset of genotyped animals (core animals) and extended that information. to non core animals by recursion. We tested several sets of core animals including 9,406 bulls with at least 1 classified daughter, 9,406 bulls and 1,052 classified dams of bulls, 9,406 bulls and 7,422 classified cows, and random samples of 5,000 to 30,000 animals. Validation reliability was assessed by the coefficient of determination from regression of daughter deviation on genomic predictions for the predicted young bulls. The reliabilities were 0.39 with 5,000 randomly chosen core animals, 0.45 with the 9,406 bulls, and 7,422 cows as core animals, and 0.44 with the remaining sets. With phenotypes truncated in 2009 and the preconditioned conjugate gradient to solve mixed model equations, the number of rounds to convergence for core animals defined by bulls was 1,343; defined by bulls and cows, 2,066; and defined by 10,000 random animals, at most 1,629. With complete phenotype data, the number of rounds decreased to 858, 1,299, and at most 1,092, respectively. Setting up G(APY)(-1) for 569,404 genotyped animals with 10,000 core animals took 1.3 h and 57 GB of memory. The validation reliability with APY reaches a plateau when the number of core animals is at least 10,000. Predictions with APY have little differences in reliability among definitions of core animals. Single-step genomic BLUP with APY is applicable to millions of genotyped animals.