An empirical evaluation of genotype imputation of ancient DNA.

An empirical evaluation of genotype imputation of ancient DNA.
复制标题

DOI:
10.1093/g3journal/jkac089
复制
发表时间:
2022-05-30
期刊:
G3 (Bethesda, Md.)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

由于对古代DNA进行测序以达到高覆盖率的能力通常受到样本质量或成本的限制,缺失基因型的插补提供了一种增加推断能力以及分析古代数据的成本效益的可能性。然而,高度的不确定性往往与古代DNA提出了几个方法上的挑战,并在这种情况下,插补方法的性能尚未得到充分探讨。为了获得进一步的见解,我们使用Beagle v4.0和来自1000个基因组项目第3阶段的参考数据对古代数据的插补进行了系统评估,调查了覆盖率,阶段性参考和研究样本量的影响。利用五个古老的个人与高覆盖率的数据,我们评估了插补数据的准确性,参考偏差,和遗传亲和力捕获的主成分分析。对于1×覆盖率的数据,我们获得了超过99%的基因型一致性水平,并且在低至0.75×的水平下具有相似的准确度和参考偏倚水平。我们的研究结果表明,使用插补数据可以是各种群体遗传分析的一个现实的选择,即使是在覆盖范围低于1×的数据。我们还表明,一个大的和不同的阶段性参考面板,以及包括低到中等覆盖率的古代个人在研究样本中可以提高插补性能,特别是对于罕见的等位基因。对遗传变异和等位基因频率的插补数据进行深入分析,可以进一步了解插补过程中出现的错误的性质,并可以为下游分析之前的后处理和验证提供实用指南。
With capabilities of sequencing ancient DNA to high coverage often limited by sample quality or cost, imputation of missing genotypes presents a possibility to increase the power of inference as well as cost-effectiveness for the analysis of ancient data. However, the high degree of uncertainty often associated with ancient DNA poses several methodological challenges, and performance of imputation methods in this context has not been fully explored. To gain further insights, we performed a systematic evaluation of imputation of ancient data using Beagle v4.0 and reference data from phase 3 of the 1000 Genomes project, investigating the effects of coverage, phased reference, and study sample size. Making use of five ancient individuals with high-coverage data available, we evaluated imputed data for accuracy, reference bias, and genetic affinities as captured by principal component analysis. We obtained genotype concordance levels of over 99% for data with 1× coverage, and similar levels of accuracy and reference bias at levels as low as 0.75×. Our findings suggest that using imputed data can be a realistic option for various population genetic analyses even for data in coverage ranges below 1×. We also show that a large and varied phased reference panel as well as the inclusion of low- to moderate-coverage ancient individuals in the study sample can increase imputation performance, particularly for rare alleles. In-depth analysis of imputed data with respect to genetic variants and allele frequencies gave further insight into the nature of errors arising during imputation, and can provide practical guidelines for postprocessing and validation prior to downstream analysis.
DOI: 10.1038/s41586-020-2378-6
发表时间: 2020-06-18
期刊: NATURE
影响因子: 64.8
作者:
Cassidy, Lara M.;Maolduin, Ros O.;Bradley, Daniel G.
通讯作者: Bradley, Daniel G.
DOI: 10.1371/journal.pgen.1008302
发表时间: 2019-07-01
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Gunther, Torsten;Nettelblad, Carl
通讯作者: Nettelblad, Carl
DOI: 10.1002/cem.750
发表时间: 2002-08-01
影响因子: 2.4
作者:
Arteaga, F;Ferrer, A
通讯作者: Ferrer, A
DOI: 10.1016/j.gde.2016.09.004
发表时间: 2016-12-01
影响因子: 4
作者:
Guenther, Torsten;Jakobsson, Mattias
通讯作者: Jakobsson, Mattias
DOI: 10.1038/ejhg.2011.10
发表时间: 2011-06-01
影响因子: 5.2
作者:
Jostins, Luke;Morley, Katherine I.;Barrett, Jeffrey C.
通讯作者: Barrett, Jeffrey C.