“Smarter” NICs for faster molecular dynamics: a case study

“Smarter” NICs for faster molecular dynamics: a case study
复制标题

“更智能”的 NIC 可实现更快的分子动力学:案例研究

DOI:
10.48550/arxiv.2204.05959
复制
发表时间:
2022
期刊:
2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
R. Vuduc
R. Vuduc
中科院分区:
--
文献类型:
--
作者:
Sara Karamati;C. Hughes;K. Hemmert;Ryan E. Grant;W. Schonbein;Scott Levy;T. Conte;Jeffrey S. Young;R. Vuduc

文献摘要

参考文献

被引文献

相似文献

这项工作评估了使用“智能”网络接口卡(SmartNIC)作为MiniMD分子动力学代理应用程序示例的计算加速器的好处。加速器是NVIDIA的BlueField-2卡,它包括一个8核Arm处理器以及少量的DRAM和存储。我们使用微基准测试和MiniMD测试了这些卡与标准英特尔服务器主机的网络和数据移动性能。在MiniMD中,我们区分了两类不同的计算,即核心计算和维护计算,它们按顺序执行。我们重构了算法和代码,以削弱这种依赖性并增加任务并行性,从而有可能增加与主机并发的BlueField-2的利用率。我们在一个由16个双插槽Intel Broadwell主机节点组成的集群上评估我们的实现,每个主机节点有一个BlueField-2。我们的结果表明,虽然BlueField-2的整体计算性能是有限的,但将它们与改进的MiniMD算法一起使用,可以在主机CPU基准上加速高达20%,而不会损失模拟精度。
This work evaluates the benefits of using a “smart” network interface card (SmartNIC) as a compute accelerator for the example of the MiniMD molecular dynamics proxy application. The accelerator is NVIDIA's BlueField-2 card, which includes an 8-core Arm processor along with a small amount of DRAM and storage. We test the networking and data movement performance of these cards compared to a standard Intel server host using microbenchmarks and MiniMD. In MiniMD, we identify two distinct classes of computation, namely core computation and maintenance computation, which are executed in sequence. We restructure the algorithm and code to weaken this dependence and increase task parallelism, thereby making it possible to increase utilization of the BlueField-2 concurrently with the host. We evaluate our implementation on a cluster consisting of 16 dual-socket Intel Broadwell host nodes with one BlueField-2 per host-node. Our results show that while the overall compute performance of BlueField-2 is limited, using them with a modified MiniMD algorithm allows for up to 20% speedup over the host CPU baseline with no loss in simulation accuracy.
使用 BlueField-2 DPU 加速现代 HPC 集群上基于 CPU 的分布式 DNN 训练
DOI: 10.1109/hoti52880.2021.00017
发表时间: 2021
期刊: 2021 IEEE Symposium on High-Performance Interconnects (HOTI
影响因子: --
作者:
Jain, Arpan;Alnaasan, Nawras;Shafi, Aamir;Subramoni, Hari;Panda, Dhabaleswar K
通讯作者: Panda, Dhabaleswar K