Sequence coverage required for accurate genotyping by sequencing in polyploid species

Sequence coverage required for accurate genotyping by sequencing in polyploid species
复制标题

多倍体物种中通过测序进行准确基因分型所需的序列覆盖度

DOI:
10.1111/1755-0998.13558
复制
发表时间:
2021-12-20
影响因子:
7.7
通讯作者:
Luo,Zewei
Luo,Zewei
中科院分区:
生物学1区
文献类型:
--
作者:
Wang,Lin;Yang,Jixuan;Luo,Zewei

文献摘要

相似文献

多倍体在真核生物的进化中发挥着重要作用,特别是对于开花植物。许多生态或农艺上重要的植物或作物物种都是多倍体,包括悬铃木(四倍体)、世界第二和第三大粮食作物小麦(六倍体)和马铃薯(四倍体)以及经济上重要的水产养殖动物,例如大西洋鲑鱼和鳟鱼。下一代测序数据能够在序列变异位点分配基因型,称为测序基因分型 (GBS)。 GBS 激发了人们对几乎所有二倍体和许多多倍体生物体的群体基因组学研究的巨大兴趣。 DNA 序列多态性是共显性的,因此可以充分了解多态性位点的潜在基因型,使 GBS 成为二倍体中的一项简单任务。然而,在多倍体物种中,序列数据通常可能无法提供信息,这使得 GBS 在多倍体中成为一项更具挑战性的任务。本文提出了新颖而严格的统计方法,用于预测确保多倍体中读段所暴露的多态性位点的准确GBS所需的序列读段数,并表明,十几个读段可以确保以95%的概率恢复任何四倍体基因型的所有组成等位基因,但需要数百个读段才能以90%的概率置信度准确揭示基因型,颠覆了文献中使用低覆盖率序列数据的GBS主张。使用四倍体马铃薯品种的 RAD-seq 数据测试了理论预测。该论文为多倍体实验学家提供了设计和进行基于序列的研究的理论指导和方法。
Polyploidy plays an important role in the evolution of eukaryotes, especially for flowering plants. Many of ecologically or agronomically important plant or crop species are polyploids, including sycamore maple (tetraploid), the world second and third largest food crops wheat (hexaploid) and potato (tetraploid) as well as economically important aquaculture animals such as Atlantic salmon and trout. The next generation sequencing data enables to allocate genotype at a sequence variant site, known as genotyping by sequencing (GBS). GBS has stimulated enormous interests in population based genomics studies in almost all diploid and many polyploid organisms. DNA sequence polymorphisms are codominant and thus fully informative about the underlying genotype at the polymorphic site, making GBS a straightforward task in diploids. However, sequence data may usually be uninformative in polyploid species, making GBS a far more challenging task in polyploids. This paper presents novel and rigorous statistical methods for predicting the number of sequence reads needed to ensure accurate GBS at a polymorphic site bared by the reads in polyploids and shows that a dozen of reads can ensure a probability of 95% to recover all constituent alleles of any tetraploid genotype but several hundreds of reads are needed to accurately uncover the genotype with probability confidence of 90%, subverting the proposition of GBS using low coverage sequence data in the literature. The theoretical prediction was tested by use of RAD‐seq data from tetraploid potato cultivars. The paper provides polyploid experimentalists with theoretical guides and methods for designing and conducting their sequence‐based studies.