High Performance Computing Framework for Tera-Scale Database Search of Mass Spectrometry Data.

High Performance Computing Framework for Tera-Scale Database Search of Mass Spectrometry Data.
复制标题

DOI:
10.1038/s43588-021-00113-z
复制
发表时间:
2021-08
期刊:
Nature computational science
影响因子:
--
通讯作者:
Saeed F
Saeed F
中科院分区:
其他
文献类型:
--
作者:
Haseeb M;Saeed F

文献摘要

被引文献

相似文献

数据库肽搜索算法推断肽从质谱(MS)数据。为了实现更大更复杂的系统生物学研究,在提高计算效率方面已经付出了大量的努力。然而,现代串行和高性能计算(HPC)算法表现出次优性能主要是由于其无效的并行设计(低资源利用率)和高开销成本。我们提出了一个HPC框架,称为HiCOPS,用于在分布式内存超级计算机上有效加速数据库肽搜索算法。HiCOPS提供了平均10倍以上的速度改进,并且比几个现有的HPC数据库搜索软件具有优越的并行性能。我们还制定了性能分析和优化的数学模型,并报告了几个关键指标的接近最佳结果,包括强规模效率、硬件利用率、负载平衡、进程间通信和I/O开销。HiCOPS中提出的核心并行设计、技术和优化是独立于搜索算法的,可以扩展到有效地加速现有和未来的算法和软件。
Database peptide search algorithms deduce peptides from mass spectrometry (MS) data. There has been substantial effort in improving their computational efficiency to achieve larger and more complex systems biology studies. However, modern serial and high-performance computing (HPC) algorithms exhibit sub-optimal performance mainly due to their ineffective parallel designs (low resource utilization), and high overhead costs. We present an HPC framework, called HiCOPS, for efficient acceleration of the database peptide search algorithms on distributed-memory supercomputers. HiCOPS provides, on average, more than 10-fold improvement in speed, and superior parallel performance over several existing HPC database search software. We also formulate a mathematical model for performance analysis and optimization, and report near-optimal results for several key metrics including strong-scale efficiency, hardware utilization, load-balance, inter-process communication and I/O overheads. The core parallel design, techniques, and optimizations presented in HiCOPS are search-algorithm independent and can be extended to efficiently accelerate the existing and future algorithms and software.