Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads

Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads
复制标题

DOI:
10.1101/gr.111120.110
复制
发表时间:
2011-06-01
期刊:
影响因子:
7
通讯作者:
Goodson, Martin
Goodson, Martin
中科院分区:
生物学1区
文献类型:
--
作者:
Lunter, Gerton;Goodson, Martin

文献摘要

被引文献

相似文献

DNA和RNA的大容量测序现在已经可以到达任何研究实验室,并迅速成为一种关键的研究工具。在许多工作流程中,由测序运行产生的每个短序列(“读取”)首先被“映射”(比对)到一个参考序列,以推断基因组位置所源自的读取,这是一项具有挑战性的任务,因为数据量很大,而且通常是大型基因组。现有读取映射软件在速度(例如,BWA、Bowtie、Eland)或灵敏度(例如,NovoAlign)方面出类拔萃,但不是两者兼而有之。此外,在存在序列变异的情况下,尤其是短插入和短缺失(INDELs)时,性能通常会恶化。在这里,我们提出了一个读映射器StamPy,它使用混合映射算法和详细的统计模型来实现速度和敏感度,特别是当读取包括序列变化时。与现有软件相比,这导致了更高的可用序列产量和更高的准确性。
High-volume sequencing of DNA and RNA is now within reach of any research laboratory and is quickly becoming established as a key research tool. In many workflows, each of the short sequences ("reads'') resulting from a sequencing run are first "mapped'' (aligned) to a reference sequence to infer the read from which the genomic location derived, a challenging task because of the high data volumes and often large genomes. Existing read mapping software excel in either speed (e. g., BWA, Bowtie, ELAND) or sensitivity (e. g., Novoalign), but not in both. In addition, performance often deteriorates in the presence of sequence variation, particularly so for short insertions and deletions (indels). Here, we present a read mapper, Stampy, which uses a hybrid mapping algorithm and a detailed statistical model to achieve both speed and sensitivity, particularly when reads include sequence variation. This results in a higher useable sequence yield and improved accuracy compared to that of existing software.