A SNP discovery method to assess variant allele probability from next-generation resequencing data

A SNP discovery method to assess variant allele probability from next-generation resequencing data
复制标题

DOI:
10.1101/gr.096388.109
复制
发表时间:
2010-02-01
期刊:
影响因子:
7
通讯作者:
Yu, Fuli
Yu, Fuli
中科院分区:
生物学1区
文献类型:
--
作者:
Shen, Yufeng;Wan, Zhengzheng;Yu, Fuli

文献摘要

被引文献

相似文献

从下一代测序(NGS)数据中准确识别遗传变异对于即时的大规模基因组研究(如1000基因组计划)至关重要,并且对于基于这些发现的进一步遗传分析至关重要。单核苷酸多态性(SNP)发现的关键挑战是区分真正的个体变异(发生在低频率)和测序错误(通常发生在高数量级的频率)。因此,了解基调用的错误概率是必要的。我们开发了Atlas-SNP2,这是一种计算工具,可以在从训练数据集学习的逻辑回归模型中检测和解释由上下文相关变量引起的系统测序错误。随后,通过贝叶斯公式估计每次替换的后验错误概率,该公式将总体测序错误概率的先验知识和估计的SNP率与给定替换的逻辑回归模型的结果相结合。估计的后验SNP概率可以用来区分真正的SNP和测序错误。验证结果表明,Atlas-SNP2的假阳性率低于10%,假阴性率接近5%或更低。
Accurate identification of genetic variants from next-generation sequencing (NGS) data is essential for immediate large-scale genomic endeavors such as the 1000 Genomes Project, and is crucial for further genetic analysis based on the discoveries. The key challenge in single nucleotide polymorphism (SNP) discovery is to distinguish true individual variants (occurring at a low frequency) from sequencing errors (often occurring at frequencies orders of magnitude higher). Therefore, knowledge of the error probabilities of base calls is essential. We have developed Atlas-SNP2, a computational tool that detects and accounts for systematic sequencing errors caused by context-related variables in a logistic regression model learned from training data sets. Subsequently, it estimates the posterior error probability for each substitution through a Bayesian formula that integrates prior knowledge of the overall sequencing error probability and the estimated SNP rate with the results from the logistic regression model for the given substitutions. The estimated posterior SNP probability can be used to distinguish true SNPs from sequencing errors. Validation results show that Atlas-SNP2 achieves a false-positive rate of lower than 10%, with an similar to 5% or lower false-negative rate.