High-performance hardware implementation of a parallel database search engine for real-time peptide mass fingerprinting

High-performance hardware implementation of a parallel database search engine for real-time peptide mass fingerprinting
复制标题

DOI:
10.1093/bioinformatics/btn216
复制
发表时间:
2008-07-01
期刊:
影响因子:
5.8
通讯作者:
Coca, Daniel
Coca, Daniel
中科院分区:
生物学3区
文献类型:
--
作者:
Bogdan, Istvan A.;Rivers, Jenny;Coca, Daniel

文献摘要

被引文献

相似文献

动机:多肽质量指纹图谱(PMF)是一种蛋白质鉴定方法,在这种方法中,蛋白质通过特定的切割程序(通常是用胰酶进行蛋白质分解)进行裂解,这些产物的质量构成了一个指纹,可以根据所有已知蛋白质的理论指纹进行搜索。在PMF的第一阶段,对原始的质谱学数据进行处理以生成多肽质量列表。在第二阶段,这个蛋白质指纹被用来搜索已知蛋白质的数据库,以寻找最佳的蛋白质匹配。尽管目前的软件解决方案通常可以在相对较短的时间内提供匹配,但可以实时找到匹配的系统可能会改变PMF的部署和呈现方式。在早些时候发表的一篇论文中,我们提出了一种原始质谱处理器的硬件设计,当在现场可编程门阵列(现场可编程门阵列)硬件中实现时,与运行在双处理器服务器上的传统软件实现相比,速度提高了近170倍。在本文中,我们提供了一个并行数据库搜索引擎的补充硬件实现,当运行在Xilinx Virtex 2 FPGA上以100 MHz的速度运行时,与运行在3.06 GHz Xeon工作站上的同等C软件例程相比,速度提高了1800倍。该设计固有的可扩展性意味着,通过将该设计部署在多个FPGA上,可以使处理速度成倍增加。数据库搜索处理器和质谱处理器运行在可重构的计算平台上,提供了完整的实时PMF蛋白质鉴定解决方案。
Motivation: Peptide mass fingerprinting (PMF) is a method for protein identification in which a protein is fragmented by a defined cleavage protocol (usually proteolysis with trypsin), and the masses of these products constitute a fingerprint that can be searched against theoretical fingerprints of all known proteins. In the first stage of PMF, the raw mass spectrometric data are processed to generate a peptide mass list. In the second stage this protein fingerprint is used to search a database of known proteins for the best protein match. Although current software solutions can typically deliver a match in a relatively short time, a system that can find a match in real time could change the way in which PMF is deployed and presented. In a paper published earlier we presented a hardware design of a raw mass spectra processor that, when implemented in Field Programmable Gate Array (FPGA) hardware, achieves almost 170-fold speed gain relative to a conventional software implementation running on a dual processor server. In this article we present a complementary hardware realization of a parallel database search engine that, when running on a Xilinx Virtex 2 FPGA at 100 MHz, delivers 1800-fold speed-up compared with an equivalent C software routine, running on a 3.06 GHz Xeon workstation. The inherent scalability of the design means that processing speed can be multiplied by deploying the design on multiple FPGAs. The database search processor and the mass spectra processor, running on a reconfigurable computing platform, provide a complete real-time PMF protein identification solution.