High quality SNP calling using Illumina data at shallow coverage

High quality SNP calling using Illumina data at shallow coverage
复制标题

DOI:
10.1093/bioinformatics/btq092
复制
发表时间:
2010-04-15
期刊:
影响因子:
5.8
通讯作者:
Jones, Steven J. M.
Jones, Steven J. M.
中科院分区:
生物学3区
文献类型:
--
作者:
Malhis, Nawar;Jones, Steven J. M.

文献摘要

被引文献

相似文献

动机:单核苷酸多态(SNPs)的检测已经成为处理第二代测序(SGS)数据的主要应用。原则上,SNPs是基于参考基因组和SGS短读取样本基因组产生的序列之间的单碱基差异而被称为SNPs。然而,这项工作远非微不足道;与测序质量和/或参考基因组属性相关的几个参数对所谓的SNPs的准确性起着至关重要的影响,特别是在覆盖范围较浅的数据中。在这项工作中,我们提出了Slider II,一种比对和SNP调用方法,展示了改进的算法方法,使更多的被调用SNP具有更低的误检率。除了常规的比对和SNP调用之外,作为一种可选功能,Slider II能够在比对和SNP调用中利用有关目标基因组的已知SNPs的信息作为先验,以增强其检测这些已知SNPs和新SNPs及其附近突变的能力。
Motivation: Detection of single nucleotide polymorphisms (SNPs) has been a major application in processing second generation sequencing (SGS) data. In principle, SNPs are called on single base differences between a reference genome and a sequence generated from SGS short reads of a sample genome. However, this exercise is far from trivial; several parameters related to sequencing quality, and/or reference genome properties, play essential effect on the accuracy of called SNPs especially at shallow coverage data. In this work, we present Slider II, an alignment and SNP calling approach that demonstrates improved algorithmic approaches enabling larger number of called SNPs with lower false positive rate. In addition to the regular alignment and SNP calling, as an optional feature, Slider II is capable of utilizing information about known SNPs of a target genome, as priors, in the alignment and SNPs calling to enhance it's capability of detecting these known SNPs and novel SNPs and mutations in their vicinity.