Estimation of Nucleotide Diversity, Disequilibrium Coefficients, and Mutation Rates from High-Coverage Genome-Sequencing Projects

Estimation of Nucleotide Diversity, Disequilibrium Coefficients, and Mutation Rates from High-Coverage Genome-Sequencing Projects
复制标题

DOI:
10.1093/molbev/msn185
复制
发表时间:
2008-11-01
影响因子:
10.7
通讯作者:
Lynch, Michael
Lynch, Michael
中科院分区:
生物学1区
文献类型:
--
作者:
Lynch, Michael

文献摘要

被引文献

相似文献

测序策略的最新进展使快速获得单个个体的高覆盖率基因组图谱成为可能,并且很快将在经济上可行,每个种群中有数百到数千个个体。虽然这些新方法为获得群体遗传参数提供了前所未有的能力,但也带来了许多挑战,最明显的是需要考虑在单个核苷酸位点上亲本等位基因的二项抽样,并消除各种序列错误来源的偏差。为了最大限度地减少这两个问题的影响,研究人员开发了一些方法,用于生成一些关键参数的几乎无偏和最小抽样方差估计,包括平均核苷酸杂合性及其在位点之间的方差,连锁不平衡分解模式与物理距离,以及自发产生突变的速率和分子谱。这些方法为有效利用人口基因组调查数据提供了一个通用平台,同时也为这类研究的优化设计提供了指导。
Recent advances in sequencing strategies have made it feasible to rapidly obtain high-coverage genomic profiles of single individuals, and soon it will be economically feasible to do so with hundreds to thousands of individuals per population. While offering unprecedented power for the acquisition of population-genetic parameters, these new methods also introduce a number of challenges, most notably the need to account for the binomial sampling of parental alleles at individual nucleotide sites and to eliminate bias from various sources of sequence errors. To minimize the effects of both problems, methods are developed for generating nearly unbiased and minimum-sampling-variance estimates of a number of key parameters, including the average nucleotide heterozygosity and its variance among sites, the pattern of decomposition of linkage disequilibrium with physical distance, and the rate and molecular spectrum of spontaneously arising mutations. These methods provide a general platform for the efficient utilization of data from population-genomic surveys, while also providing guidance for the optimal design of such studies.