SHRiMP: accurate mapping of short color-space reads.

SHRiMP: accurate mapping of short color-space reads.
复制标题

DOI:
10.1371/journal.pcbi.1000386
复制
发表时间:
2009-05
影响因子:
4.3
通讯作者:
Brudno M
Brudno M
中科院分区:
生物学2区
文献类型:
--
作者:
Rumble SM;Lacroute P;Dalca AV;Fiume M;Sidow A;Brudno M

文献摘要

参考文献

被引文献

相似文献

下一代测序技术的发展,能够在一次运行中测序数亿个短读段(每个25-70 bp),为非模式物种的群体基因组研究打开了大门。在本文中,我们提出了SHRiMP -短读映射包:一套算法和方法,映射短读到基因组,即使在存在大量的多态性。我们的方法是基于一个快速读取映射技术,单独的彻底对齐方法,定期字母空间以及AB SOLiD(颜色空间)读取,和假阳性命中的统计模型。我们使用SHRiMP将来自新测序的玻璃海鞘个体的读段映射到参考基因组。我们证明,SHRiMP可以准确地映射读段到这个高度多态性的基因组,同时确认C.萨维尼在这第二个人。SHRiMP可在http://compbio.cs.toronto.edu/shrimp上免费获得。下一代测序(NGS)技术正在彻底改变生物学家获取和分析基因组数据的方式。NGS机器,如Illumina/Solexa和AB SOLiD,能够比以前的方法便宜200倍地进行基因组测序。NGS技术的主要应用领域之一是发现给定物种内的基因组变异。发现这种变异的第一步是将从供体个体测序的读数映射到已知(“参考”)基因组。参考和读数之间的差异指示多态性或测序错误。自从引入NGS技术以来,已经设计了许多方法用于将读段映射到参考基因组。然而,这些算法往往牺牲灵敏度的快速运行时间。虽然它们在映射来自表现出低多态性率的生物体的读段方面是成功的,但它们在映射来自高度多态性生物体的读段方面表现不佳。我们提出了一种新的读取映射方法,SHRiMP,可以处理更大量的多态性。使用玻璃海鞘作为我们的目标生物,我们证明了我们的方法发现显着更多的变化比其他方法。此外,我们开发了经典比对算法的颜色空间扩展,使我们能够映射颜色空间或“二碱基”,由AB SOLiD测序仪生成的读数。
The development of Next Generation Sequencing technologies, capable of sequencing hundreds of millions of short reads (25–70 bp each) in a single run, is opening the door to population genomic studies of non-model species. In this paper we present SHRiMP - the SHort Read Mapping Package: a set of algorithms and methods to map short reads to a genome, even in the presence of a large amount of polymorphism. Our method is based upon a fast read mapping technique, separate thorough alignment methods for regular letter-space as well as AB SOLiD (color-space) reads, and a statistical model for false positive hits. We use SHRiMP to map reads from a newly sequenced Ciona savignyi individual to the reference genome. We demonstrate that SHRiMP can accurately map reads to this highly polymorphic genome, while confirming high heterozygosity of C. savignyi in this second individual. SHRiMP is freely available at http://compbio.cs.toronto.edu/shrimp. Next Generation Sequencing (NGS) technologies are revolutionizing the way biologists acquire and analyze genomic data. NGS machines, such as Illumina/Solexa and AB SOLiD, are able to sequence genomes more cheaply by 200-fold than previous methods. One of the main application areas of NGS technologies is the discovery of genomic variation within a given species. The first step in discovering this variation is the mapping of reads sequenced from a donor individual to a known (“reference”) genome. Differences between the reference and the reads are indicative either of polymorphisms, or of sequencing errors. Since the introduction of NGS technologies, many methods have been devised for mapping reads to reference genomes. However, these algorithms often sacrifice sensitivity for fast running time. While they are successful at mapping reads from organisms that exhibit low polymorphism rates, they do not perform well at mapping reads from highly polymorphic organisms. We present a novel read mapping method, SHRiMP, that can handle much greater amounts of polymorphism. Using Ciona savignyi as our target organism, we demonstrate that our method discovers significantly more variation than other methods. Additionally, we develop color-space extensions to classical alignment algorithms, allowing us to map color-space, or “dibase”, reads generated by AB SOLiD sequencers.
DOI: 10.1093/bioinformatics/16.8.699
发表时间: 2000-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rognes, T;Seeberg, E
通讯作者: Seeberg, E
DOI: 10.1089/10665270252935430
发表时间: 2002-01-01
影响因子: 1.7
作者:
Buhler, J;Tompa, M
通讯作者: Tompa, M
亚洲个体的二倍体基因组序列
DOI: 10.1038/nature07484
发表时间: 2008-11-06
期刊: NATURE
影响因子: 64.8
作者:
Wang, Jun;Wang, Wei;Li, Ruiqiang;Li, Yingrui;Tian, Geng;Goodman, Laurie;Fan, Wei;Zhang, Junqing;Li, Jun;Zhang, Juanbin;Guo, Yiran;Feng, Binxiao;Li, Heng;Lu, Yao;Fang, Xiaodong;Liang, Huiqing;Du, Zhenglin;Li, Dong;Zhao, Yiqing;Hu, Yujie;Yang, Zhenzhen;Zheng, Hancheng;Hellmann, Ines;Inouye, Michael;Pool, John;Yi, Xin;Zhao, Jing;Duan, Jinjie;Zhou, Yan;Qin, Junjie;Ma, Lijia;Li, Guoqing;Yang, Zhentao;Zhang, Guojie;Yang, Bin;Yu, Chang;Liang, Fang;Li, Wenjie;Li, Shaochuan;Li, Dawei;Ni, Peixiang;Ruan, Jue;Li, Qibin;Zhu, Hongmei;Liu, Dongyuan;Lu, Zhike;Li, Ning;Guo, Guangwu;Zhang, Jianguo;Ye, Jia;Fang, Lin;Hao, Qin;Chen, Quan;Liang, Yu;Su, Yeyang;San, A.;Ping, Cuo;Yang, Shuang;Chen, Fang;Li, Li;Zhou, Ke;Zheng, Hongkun;Ren, Yuanyuan;Yang, Ling;Gao, Yang;Yang, Guohua;Li, Zhuo;Feng, Xiaoli;Kristiansen, Karsten;Wong, Gane Ka-Shu;Nielsen, Rasmus;Durbin, Richard;Bolund, Lars;Zhang, Xiuqing;Li, Songgang;Yang, Huanming;Wang, Jian
通讯作者: Wang, Jian
DOI: 10.1038/nature07485
发表时间: 2008-11-06
期刊: NATURE
影响因子: 64.8
作者:
Ley, Timothy J.;Mardis, Elaine R.;Ding, Li;Fulton, Bob;McLellan, Michael D.;Chen, Ken;Dooling, David;Dunford-Shore, Brian H.;McGrath, Sean;Hickenbotham, Matthew;Cook, Lisa;Abbott, Rachel;Larson, David E.;Koboldt, Dan C.;Pohl, Craig;Smith, Scott;Hawkins, Amy;Abbott, Scott;Locke, Devin;Hillier, LaDeana W.;Miner, Tracie;Fulton, Lucinda;Magrini, Vincent;Wylie, Todd;Glasscock, Jarret;Conyers, Joshua;Sander, Nathan;Shi, Xiaoqi;Osborne, John R.;Minx, Patrick;Gordon, David;Chinwalla, Asif;Zhao, Yu;Ries, Rhonda E.;Payton, Jacqueline E.;Westervelt, Peter;Tomasson, Michael H.;Watson, Mark;Baty, Jack;Ivanovich, Jennifer;Heath, Sharon;Shannon, William D.;Nagarajan, Rakesh;Walter, Matthew J.;Link, Daniel C.;Graubert, Timothy A.;DiPersio, John F.;Wilson, Richard K.
通讯作者: Wilson, Richard K.
DOI: 10.1142/s0219720004000661
发表时间: 2004-09-01
影响因子: 1
作者:
Li, Ming;Ma, Bin;Tromp, John
通讯作者: Tromp, John