Fast imputation using medium or low-coverage sequence data.

Fast imputation using medium or low-coverage sequence data.
复制标题

DOI:
10.1186/s12863-015-0243-7
复制
发表时间:
2015-07-14
期刊:
影响因子:
2.9
通讯作者:
O'Connell JR
O'Connell JR
中科院分区:
生物学3区
文献类型:
--
作者:
VanRaden PM;Sun C;O'Connell JR

文献摘要

被引文献

相似文献

通过结合不同读取深度的全基因组序列数据和不同密度的阵列基因型,准确的基因型插补可以大大降低成本并增加效益。对于大量群体,有效的策略选择最有可能形成每种基因型的两个单倍型,并在处理每个个体的序列时根据这两个单倍型内的先验概率更新后等位基因概率。与先调用或计算基因型概率然后插补相比,直接使用等位基因读数计数可以提高插补准确性并减少计算量。新算法在 findhap(版本 4)软件中实施,并使用模拟牛和实际人类序列数据以及参考群体大小、序列读取深度和错误率的不同组合进行测试。对于测序个体的直接研究可能需要 ≥8× 的读数深度,但对于给定的总成本,以 2× 至 4× 的读数深度对更多个体进行测序可以从阵列基因型中获得更准确的插补。如果参考个体同时具有低覆盖率序列和高密度 (HD) 微阵列数据,则插补准确性会进一步提高,并且即使读取错误率为 16%,插补准确性仍保持较高水平。在读取深度≤4×的情况下,findhap(版本4)的准确率高于Beagle(版本4); findhap 的计算时间比 Beagle 快 400 倍。对于 10,000 个已测序个体以及 250 个具有 HD 阵列基因型的个体来测试插补,findhap 使用了 7 小时、10 个处理器和 50 GB 内存来处理一条染色体上的 100 万个位点。计算时间与种群规模成正比,但与变体数量成正比。通过更新每个个体的两个单倍型内的等位基因概率,在 findhap 中可以非常有效地同时从低覆盖率序列数据调用基因型并从不同密度的阵列基因型进行插补。模拟牛和实际人类基因组均简化为低覆盖率序列和 HD 微阵列数据,基因型识别和插补的准确性很高。更有效的插补使遗传学家能够定位和测试来自更多个体的更多 DNA 变异的影响,并将其纳入未来的预测和选择中。
Accurate genotype imputation can greatly reduce costs and increase benefits by combining whole-genome sequence data of varying read depth and array genotypes of varying densities. For large populations, an efficient strategy chooses the two haplotypes most likely to form each genotype and updates posterior allele probabilities from prior probabilities within those two haplotypes as each individual’s sequence is processed. Directly using allele read counts can improve imputation accuracy and reduce computation compared with calling or computing genotype probabilities first and then imputing. A new algorithm was implemented in findhap (version 4) software and tested using simulated bovine and actual human sequence data with different combinations of reference population size, sequence read depth and error rate. Read depths of ≥8× may be desired for direct investigation of sequenced individuals, but for a given total cost, sequencing more individuals at read depths of 2× to 4× gave more accurate imputation from array genotypes. Imputation accuracy improved further if reference individuals had both low-coverage sequence and high-density (HD) microarray data, and remained high even with a read error rate of 16 %. With read depths of ≤4×, findhap (version 4) had higher accuracy than Beagle (version 4); computing time was up to 400 times faster with findhap than with Beagle. For 10,000 sequenced individuals plus 250 with HD array genotypes to test imputation, findhap used 7 hours, 10 processors and 50 GB of memory for 1 million loci on one chromosome. Computing times increased in proportion to population size but less than proportional to number of variants. Simultaneous genotype calling from low-coverage sequence data and imputation from array genotypes of various densities is done very efficiently within findhap by updating allele probabilities within the two haplotypes for each individual. Accuracy of genotype calling and imputation were high with both simulated bovine and actual human genomes reduced to low-coverage sequence and HD microarray data. More efficient imputation allows geneticists to locate and test effects of more DNA variants from more individuals and to include those in future prediction and selection.