Hardware Acceleration of Tensor-Structured Multilevel Ewald Summation Method on MDGRAPE-4A, a Special-Purpose Computer System for Molecular Dynamics Simulations

Hardware Acceleration of Tensor-Structured Multilevel Ewald Summation Method on MDGRAPE-4A, a Special-Purpose Computer System for Molecular Dynamics Simulations
复制标题

DOI:
10.1145/3458817.3476190
复制
发表时间:
2021-11
期刊:
SC21: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
G. Morimoto;Yohei M. Koyama;Hao Zhang;T. Komatsu;Y. Ohno;Keigo Nishida;Itta Ohmura;H. Koyama;M. Taiji
G. Morimoto;Yohei M. Koyama;Hao Zhang;T. Komatsu;Y. Ohno;Keigo Nishida;Itta Ohmura;H. Koyama;M. Taiji
中科院分区:
其他
文献类型:
--
作者:
G. Morimoto;Yohei M. Koyama;Hao Zhang;T. Komatsu;Y. Ohno;Keigo Nishida;Itta Ohmura;H. Koyama;M. Taiji

文献摘要

相似文献

我们开发了mdgraph - 4a,一种用于分子动力学模拟的专用计算机系统,由512个定制的单片系统lsi节点组成,具有专用处理器内核和互连,旨在实现生物分子模拟的强大可扩展性。为了减少库仑相互作用评估所需的全局通信,我们进行了mdgraph - 4a和新算法(张量结构多层埃瓦尔德求和法(TME))的联合设计,该算法在定制LSI电路上为粒子网格操作和三维环面网络上的网格-网格可分离卷积生成硬件模块。我们通过在FPGA上使用3D fft以及基于FPGA的八叉树网络来收集网格电荷,实现了顶层网格电位的卷积。库仑远程部分的运行时间为50 $\mu\mathrm{s}$,与短程部分的运行时间基本重合,额外的开销约为10 $\mu\mathrm{s}/\text{step}$,性能损失仅为5%。
We developed MDGRAPE-4A, a special-purpose computer system for molecular dynamics simulations, consisting of 512 nodes of custom system-on-a-chip LSIs with dedicated processor cores and interconnects designed to achieve strong scalability for biomolecular simulations. To reduce the global communications required for the evaluation of Coulomb interactions, we conducted a co-design of the MDGRAPE-4A and the novel algorithm, tensor-structured multilevel Ewald summation method (TME), which produced hardware modules on the custom LSI circuit for particle-grid operations and for grid-grid separable convolutions on a 3D torus network. We implemented the convolution for the top-level grid potentials by using 3D FFTs on an FPGA, along with an FPGA-based octree network to gather grid charges. The elapsed time for the long-range part of Coulomb is 50 $\mu\mathrm{s}$, which can mostly overlap with those for the short-range part, and the additional cost is approximately 10 $\mu\mathrm{s}/\text{step}$, which is only a 5% performance loss.