Population genomics based on low coverage sequencing: how low should we go?

Population genomics based on low coverage sequencing: how low should we go?
复制标题

DOI:
10.1111/mec.12105
复制
发表时间:
2013-06-01
期刊:
影响因子:
4.9
通讯作者:
Gompert, Zachariah
Gompert, Zachariah
中科院分区:
生物学1区
文献类型:
--
作者:
Buerkle, C. Alex;Gompert, Zachariah

文献摘要

被引文献

相似文献

分子生态学的研究现在往往基于大量的DNA序列读取。考虑到DNA测序的时间和财政预算,问题就出现了,如何在三个维度中分配有限数量的序列读数:(i)重复测序个体核苷酸位置并对个体的真实基因型获得高置信度,(ii)从群体中取样更多的个体,(iii)取样更大的基因组部分。先不考虑采样基因组的哪个部分的问题,我们分析了重复测序相同的核苷酸位置(覆盖深度)和样本中个体数量之间的权衡。我们回顾了等位基因频率的简单贝叶斯模型,并利用这些模型来分析如何获得最大的群体遗传参数信息。模型表明,以牺牲每个核苷酸位置的覆盖深度为代价,对更多的个体进行采样,可以提供更多关于总体参数的信息。将测序工作最大限度地分配给个体,并获得每个位点和个体大约一个读数(1倍覆盖率),从而获得关于种群的最多信息。一些分析需要个体的遗传参数,在这种情况下,贝叶斯种群模型也支持从较低覆盖序列数据推断,而不是简单的似然模型。低覆盖率测序不仅足以支持推断,而且设计利用低覆盖率的研究是最佳的,因为它们将基于基因组中更多的个体或位点产生高度准确和精确的参数估计。
Research in molecular ecology is now often based on large numbers of DNA sequence reads. Given a time and financial budget for DNA sequencing, the question arises as to how to allocate the finite number of sequence reads among three dimensions: (i) sequencing individual nucleotide positions repeatedly and achieving high confidence in the true genotype of individuals, (ii) sampling larger numbers of individuals from a population, and (iii) sampling a larger fraction of the genome. Leaving aside the question of what fraction of the genome to sample, we analyze the trade-off between repeatedly sequencing the same nucleotide position (coverage depth) and the number of individuals in the sample. We review simple Bayesian models for allele frequencies and utilize these in the analysis of how to obtain maximal information about population genetic parameters. The models indicate that sampling larger numbers of individuals, at the expense of coverage depth per nucleotide position, provides more information about population parameters. Dividing the sequencing effort maximally among individuals and obtaining approximately one read per locus and individual (1xcoverage) yields the most information about a population. Some analyses require genetic parameters for individuals, in which case Bayesian population models also support inference from lower coverage sequence data than are required for simple likelihood models. Low coverage sequencing is not only sufficient to support inference, but it is optimal to design studies to utilize low coverage because they will yield highly accurate and precise parameter estimates based on more individuals or sites in the genome.