SiNPle: Fast and Sensitive Variant Calling for Deep Sequencing Data

SiNPle: Fast and Sensitive Variant Calling for Deep Sequencing Data
复制标题

DOI:
10.3390/genes10080561
复制
发表时间:
2019-08-01
期刊:
影响因子:
3.5
通讯作者:
Ribeca, Paolo
Ribeca, Paolo
中科院分区:
生物学3区
文献类型:
--
作者:
Ferretti, Luca;Tennakoon, Chandana;Ribeca, Paolo

文献摘要

被引文献

相似文献

目前的高通量测序技术可以生成序列数据,并以非常高的覆盖率提供关于样品遗传组成的信息。深度测序方法能够检测异质样品中的罕见变体,例如病毒准种,但也具有放大测序错误和伪影的不期望的效果。区分真实的变体与这种噪声并不简单。可以处理合并样本的变异识别者在极高的读取深度可能会遇到麻烦,而在较低的深度,灵敏度往往会牺牲特异性。在本文中,我们提出了SiNPle(简化推理的新的多态性从大覆盖率),一个快速,有效的软件变异调用。SiNPle基于简化的贝叶斯方法来计算变异不是由测序错误或PCR伪影产生的后验概率。贝叶斯模型考虑了个体碱基质量及其分布、测序和PCR阶段的基线错误率、变异频率的先验分布及其链化。我们的方法导致一个近似的,但非常快速的计算后验概率,即使是非常高的覆盖率的数据,因为后验分布的表达式是一个简单的分析公式,在基因组中的每个站点出现的变异的汇总统计。这些统计数据可用于根据所需的灵敏度水平过滤出推定的SNP和插入缺失。我们在几个模拟和真实的病毒数据集上测试了SiNPle,以证明它比现有方法更快,更灵敏。SiNPle的源代码可以免费下载和编译,或者作为Conda/Bioconda包。
Current high-throughput sequencing technologies can generate sequence data and provide information on the genetic composition of samples at very high coverage. Deep sequencing approaches enable the detection of rare variants in heterogeneous samples, such as viral quasi-species, but also have the undesired effect of amplifying sequencing errors and artefacts. Distinguishing real variants from such noise is not straightforward. Variant callers that can handle pooled samples can be in trouble at extremely high read depths, while at lower depths sensitivity is often sacrificed to specificity. In this paper, we propose SiNPle (Simplified Inference of Novel Polymorphisms from Large coveragE), a fast and effective software for variant calling. SiNPle is based on a simplified Bayesian approach to compute the posterior probability that a variant is not generated by sequencing errors or PCR artefacts. The Bayesian model takes into consideration individual base qualities as well as their distribution, the baseline error rates during both the sequencing and the PCR stage, the prior distribution of variant frequencies and their strandedness. Our approach leads to an approximate but extremely fast computation of posterior probabilities even for very high coverage data, since the expression for the posterior distribution is a simple analytical formula in terms of summary statistics for the variants appearing at each site in the genome. These statistics can be used to filter out putative SNPs and indels according to the required level of sensitivity. We tested SiNPle on several simulated and real-life viral datasets to show that it is faster and more sensitive than existing methods. The source code for SiNPle is freely available to download and compile, or as a Conda/Bioconda package.