Molecular Dynamics Range-Limited Force Evaluation Optimized for FPGAs

Molecular Dynamics Range-Limited Force Evaluation Optimized for FPGAs
复制标题

针对 FPGA 优化的分子动力学范围限制力评估

DOI:
10.1109/asap.2019.00016
复制
发表时间:
2019
期刊:
2019 IEEE 30th International Conference on Application-specific Systems, Architectures and Processors (ASAP)
影响因子:
--
通讯作者:
M. Herbordt
M. Herbordt
中科院分区:
--
文献类型:
--
作者:
Chen Yang;Tong Geng;Tianqi Wang;Charles Lin;Jiayi Sheng;Vipin Sachdeva;W. Sherman;M. Herbordt

文献摘要

被引文献

相似文献

FPGA分子动力学从2004 - 2010年进行了大量研究。由于该时代的芯片资源有限,以及包括分子动力学模拟(MD)的任务的固有多样性和复杂性,依靠主机或嵌入式处理器来组织和预处理输入和输出数据的FPGA加速器。这引入了模拟迭代之间的数据流动,并且随着技术的进步,性能非常有限。当前一代FPGA不仅配备了丰富的片上资源,而且还具有用于浮点操作的硬件支持;这些进步为在单个设备上创建独立的MD模拟系统提供了机会。在本文中,我们演示了基于范围限制力的系统,该系统包括典型的MD模拟中90%的拖鞋。它具有在线粒子对生成,数百个力评估管道,运动更新和粒子数据迁移。我们整合到OpenMM中,发现对于代表性数据集(具有20K原子的液体氩),我们可以使用单个FPGA实现1.4US/天的模拟吞吐量,是可比一代GPU的两倍以上。这里介绍的大部分工作探讨了针对现代FPGA量身定制的独立MD范围限制力评估系统的设计,而无需与任何芯片外设备进行数据交换。主要贡献是新功能的设计,将这些特征耦合到集成系统中的方法,尤其是对粒子/细胞之间最可能映射的分析,片上记忆(BRAMS)和片上的分析计算单元(管道)。
FPGA Molecular Dynamics was much studied from 2004-2010. Due to limited chip resources of that era, and the inherent variety and complexity of tasks comprising Molecular Dynamics simulations (MD), those FPGA accelerators relied on host or embedded processors to organize and pre-process input and output data. This introduced long latency for data movement between simulation iterations and, as technology advanced, drastically limited performance. Current generation FPGAs are equipped not only with abundant on-chip resources, but also have hardware support for floating point operations; these advances provide an opportunity for creating self-contained MD simulation systems on a single device. In this paper, we demonstrate such a system based on the range-limited force, which comprises 90% of the flops in a typical MD simulation. It features online particle-pair generation, hundreds of force evaluation pipelines, motion update, and particle data migration. We integrate into OpenMM and find that, for a representative dataset (liquid argon with 20K atoms), we can achieve a simulation throughput of 1.4us/day with a single FPGA, more than twice the performance of a comparable generation GPU. The bulk of the work presented here explores the design of an independent MD range-limited force evaluation system tailored for modern FPGAs without data exchange with any off-chip devices. The primary contributions are the designs of the new features, the methods for coupling those features into an integrated system, and, especially, the analysis of the most likely mappings among particles/cells, on-chip memories (BRAMs), and on-chip compute units (pipelines).