Accurate Genotype Imputation in Multiparental Populations from Low-Coverage Sequence

Accurate Genotype Imputation in Multiparental Populations from Low-Coverage Sequence
复制标题

DOI:
10.1534/genetics.118.300885
复制
发表时间:
2018-09-01
期刊:
影响因子:
3.3
通讯作者:
van Eeuwijk, Fred A.
van Eeuwijk, Fred A.
中科院分区:
生物学2区
文献类型:
--
作者:
Zheng, Chaozhi;Boer, Martin P.;van Eeuwijk, Fred A.

文献摘要

被引文献

相似文献

最近,许多不同类型的多亲群体被创造出来,以增加QTL定位中的遗传多样性和分辨率。低覆盖率、测序基因分型(GBS)技术在这些人群中已成为一种成本效益高的工具,尽管后代和创始人中存在大量缺失数据。在这项工作中,我们从低覆盖率的GBS数据中提出了一个在这些实验杂交中进行基因归属的通用统计框架。推广了以前发展的用于计算后代DNA祖先起源的隐马尔可夫模型,提出了一种不需要父母数据的适用于双亲和多亲群体的推算算法。我们的补偿算法允许双亲和后代的杂合性以及观察到的基因类型的误差校正。此外,我们的方法可以结合来自测序读数的归因和基因调用,它也适用于来自SNP阵列数据的被调用的基因类型。我们在四个不同类型的群体:F-2群体、先进杂交重组自交系、多亲先进世代杂交群体和异花授粉群体中,通过模拟和真实数据集对我们的算法进行了评估。由于我们的方法有效地利用了标记数据和群体设计信息,与以前的方法的比较表明,即使在很低的(13)测序深度下,我们的推算也是准确的,此外还具有准确的基因分相和错误检测。
Many different types of multiparental populations have recently been produced to increase genetic diversity and resolution in QTL mapping. Low-coverage, genotyping-by-sequencing (GBS) technology has become a cost-effective tool in these populations, despite large amounts of missing data in offspring and founders. In this work, we present a general statistical framework for genotype imputation in such experimental crosses from low-coverage GBS data. Generalizing a previously developed hidden Markov model for calculating ancestral origins of offspring DNA, we present an imputation algorithm that does not require parental data and that is applicable to bi- and multiparental populations. Our imputation algorithm allows heterozygosity of parents and offspring as well as error correction in observed genotypes. Further, our approach can combine imputation and genotype calling from sequencing reads, and it also applies to called genotypes from SNP array data. We evaluate our imputation algorithm by simulated and real data sets in four different types of populations: the F-2, the advanced intercross recombinant inbred lines, the multiparent advanced generation intercross, and the cross-pollinated population. Because our approach uses marker data and population design information efficiently, the comparisons with previous approaches show that our imputation is accurate at even very low (, 13) sequencing depth, in addition to having accurate genotype phasing and error detection.