A hybrid method for the imputation of genomic data in livestock populations.

A hybrid method for the imputation of genomic data in livestock populations.
复制标题

DOI:
10.1186/s12711-017-0300-y
复制
发表时间:
2017-03-03
期刊:
Genetics, selection, evolution : GSE
影响因子:
--
通讯作者:
Hickey JM
Hickey JM
中科院分区:
其他
文献类型:
--
作者:
Antolín R;Nettelblad C;Gorjanc G;Money D;Hickey JM

文献摘要

被引文献

相似文献

本文描述了一种结合启发式和隐马尔可夫模型(HMM)的方法,以准确地填补牲畜数据集的缺失基因型。育种计划中的基因组选择需要对许多个体进行高密度的基因分型,这使得经济地产生这些信息的算法至关重要。有两类常见的估算方法,启发式方法和概率方法,后者主要是基于隐马尔可夫模型。启发式方法是强大的,但不能归因于在启发式规则的阈值不满足,或系谱是不一致的区域的标记。隐马尔可夫模型是一种概率方法,通常不需要特定的家族结构或谱系信息,这使得它们非常灵活,但它们在计算上昂贵,并且在某些情况下不太准确。我们实现了一种新的混合插补方法,结合启发式和HMM方法,AlphaImpute和MaCH,并比较了三种方法的计算时间和插补精度。AlphaImpute是最快的,其次是混合方法,然后是HMM。混合方法和隐马尔可夫模型的计算时间与隐马尔可夫模型中使用的迭代次数呈线性增加,然而,混合方法的计算时间几乎呈线性增加,而隐马尔可夫模型的计算时间与模板单倍型的数量呈二次方增加。混合法是最准确的插补方法的低密度面板时,系谱信息缺失,特别是如果次要等位基因频率也很低。混合方法和HMM的准确性随着模板单倍型的数量而增加。所有三种方法的插补准确性随着低密度板的标记密度而增加。排除系谱信息降低了混合方法和AlphaImpute的插补准确度。最后,三种方法的插补精度随着次要等位基因频率的降低而降低。混合启发式和概率插补方法能够插补群体中所有个体的所有标记,如HMM。混合方法通常更准确,并且永远不会比纯粹的启发式方法或纯粹的概率方法准确得多,并且比标准的概率方法更快。本文的在线版本(doi:10.1186/s12711-017-0300-y)包含补充材料,可供授权用户使用。
This paper describes a combined heuristic and hidden Markov model (HMM) method to accurately impute missing genotypes in livestock datasets. Genomic selection in breeding programs requires high-density genotyping of many individuals, making algorithms that economically generate this information crucial. There are two common classes of imputation methods, heuristic methods and probabilistic methods, the latter being largely based on hidden Markov models. Heuristic methods are robust, but fail to impute markers in regions where the thresholds of heuristic rules are not met, or the pedigree is inconsistent. Hidden Markov models are probabilistic methods which typically do not require specific family structures or pedigree information, making them very flexible, but they are computationally expensive and, in some cases, less accurate. We implemented a new hybrid imputation method that combined heuristic and HMM methods, AlphaImpute and MaCH, and compared the computation time and imputation accuracy of the three methods. AlphaImpute was the fastest, followed by the hybrid method and then the HMM. The computation time of the hybrid method and the HMM increased linearly with the number of iterations used in the hidden Markov model, however, the computation time of the hybrid method increased almost linearly and that of the HMM quadratically with the number of template haplotypes. The hybrid method was the most accurate imputation method for low-density panels when pedigree information was missing, especially if minor allele frequency was also low. The accuracy of the hybrid method and the HMM increased with the number of template haplotypes. The imputation accuracy of all three methods increased with the marker density of the low-density panels. Excluding the pedigree information reduced imputation accuracy for the hybrid method and AlphaImpute. Finally, the imputation accuracy of the three methods decreased with decreasing minor allele frequency. The hybrid heuristic and probabilistic imputation method is able to impute all markers for all individuals in a population, as the HMM. The hybrid method is usually more accurate and never significantly less accurate than a purely heuristic method or a purely probabilistic method and is faster than a standard probabilistic method. The online version of this article (doi:10.1186/s12711-017-0300-y) contains supplementary material, which is available to authorized users.