Compressive genomics for protein databases.

Compressive genomics for protein databases.
复制标题

DOI:
10.1093/bioinformatics/btt214
复制
发表时间:
2013-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Berger B
Berger B
中科院分区:
其他
文献类型:
--
作者:
Daniels NM;Gallant A;Peng J;Cowen LJ;Baym M;Berger B

文献摘要

参考文献

被引文献

相似文献

动机:蛋白质序列数据库的指数增长日益使搜索同源物的基本问题成为计算瓶颈。然而,独特数据量的增长速度却没有那么快。我们可以利用这一事实来大大加速同源搜索。流行的 PSI/DELTA-BLAST 系列工具中的程序加速不仅会直接加速同源性搜索,而且还会加速主要通过这些工具与大型蛋白质数据库交互的其他当前程序的大量集合。结果:我们推出了一套由压缩加速蛋白 BLAST (CaBLASTP) 提供支持的同源性搜索工具,其速度明显快于所有已知的最先进工具,包括 HHblits、DELTA-BLAST 和 PSI-BLAST,并且相当准确。此外,我们的工具的实施方式允许直接替换现有的分析管道。关键思想是我们引入了一种基于局部相似性的压缩方案,该方案允许我们直接对压缩数据进行操作。重要的是,CaBLASTP 的运行时间在独特数据量上几乎呈线性缩放,而当前的 BLASTP 变体则在搜索的完整蛋白质数据库的大小上呈线性缩放。我们的压缩算法将加快许多任务的速度,例如蛋白质结构预测和直系映射,这些任务严重依赖同源搜索。可用性:CaBLASTP 可根据 GNU 公共许可证在 http://cablastp.csail.mit.edu/ 上获取 联系方式:bab@mit.edu
Motivation: The exponential growth of protein sequence databases has increasingly made the fundamental question of searching for homologs a computational bottleneck. The amount of unique data, however, is not growing nearly as fast; we can exploit this fact to greatly accelerate homology search. Acceleration of programs in the popular PSI/DELTA-BLAST family of tools will not only speed-up homology search directly but also the huge collection of other current programs that primarily interact with large protein databases via precisely these tools. Results: We introduce a suite of homology search tools, powered by compressively accelerated protein BLAST (CaBLASTP), which are significantly faster than and comparably accurate with all known state-of-the-art tools, including HHblits, DELTA-BLAST and PSI-BLAST. Further, our tools are implemented in a manner that allows direct substitution into existing analysis pipelines. The key idea is that we introduce a local similarity-based compression scheme that allows us to operate directly on the compressed data. Importantly, CaBLASTP’s runtime scales almost linearly in the amount of unique data, as opposed to current BLASTP variants, which scale linearly in the size of the full protein database being searched. Our compressive algorithms will speed-up many tasks, such as protein structure prediction and orthology mapping, which rely heavily on homology search. Availability: CaBLASTP is available under the GNU Public License at http://cablastp.csail.mit.edu/ Contact: bab@mit.edu
DOI: 10.1093/bioinformatics/bts110
发表时间: 2012-05-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Daniels NM;Hosur R;Berger B;Cowen LJ
通讯作者: Cowen LJ
DOI: 10.1093/bioinformatics/btp265
发表时间: 2009-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Kumar A;Cowen L
通讯作者: Cowen L
DOI: 10.1371/journal.pcbi.1000779
发表时间: 2010-05-27
影响因子: 4.3
作者:
Huttenhower C;Hofmann O
通讯作者: Hofmann O
DOI: 10.1002/prot.21770
发表时间: 2008-05-01
影响因子: 2.9
作者:
Kosloff, Mickey;Kolodny, Rachel
通讯作者: Kolodny, Rachel
DOI: 10.1093/bioinformatics/18.12.1696
发表时间: 2002-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Chen, X;Li, M;Tromp, J
通讯作者: Tromp, J