Accounting for bias from sequencing error in population genetic estimates

Accounting for bias from sequencing error in population genetic estimates
复制标题

DOI:
10.1093/molbev/msm239
复制
发表时间:
2008-01-01
影响因子:
10.7
通讯作者:
Slatkin, Montgomery
Slatkin, Montgomery
中科院分区:
生物学1区
文献类型:
--
作者:
Johnson, Philip L. F.;Slatkin, Montgomery

文献摘要

被引文献

相似文献

测序误差对一般使用低覆盖率序列,特别是单次读取的群体遗传分析提出了重大挑战。当多态性(信号)的水平相对于误差(噪声)的数量较低时,参数估计中的偏差就会变得严重。选择任意的质量分数截止值会产生有偏差的估计,特别是对于具有不同质量分数分布的较新的非sanger测序技术。我们提出了一个经验法则来判断一个给定的阈值何时会导致显著的偏差,并提出了减少偏差的替代方法。
Sequencing error presents a significant challenge to population genetic analyses using low-coverage sequence in general and single-pass reads in particular. Bias in parameter estimates becomes severe when the level of polymorphism (signal) is low relative to the amount of error (noise). Choosing an arbitrary quality score cutoff yields biased estimates, particularly with newer, non-Sanger sequencing technologies that have different quality score distributions. We propose a rule of thumb to judge when a given threshold will lead to significant bias and suggest alternative approaches that reduce bias.