Toward FPGA-Based HPC: Advancing Interconnect Technologies

Toward FPGA-Based HPC: Advancing Interconnect Technologies
复制标题

DOI:
10.1109/mm.2019.2950655
复制
发表时间:
2020-01-01
期刊:
影响因子:
3.6
通讯作者:
Goodacre, John
Goodacre, John
中科院分区:
计算机科学3区
文献类型:
--
作者:
Lant, Joshua;Navaridas, Javier;Goodacre, John

文献摘要

被引文献

相似文献

高性能计算架构师目前面临着来自日益严格的功率限制和不断变化的工作负载特征的无数挑战。在本文中,我们讨论了高性能计算系统中fpga的现状。最近的技术进步表明,它们在进入高性能计算市场方面处于有利地位。然而,仍有一些研究问题需要克服;我们解决了系统架构和互连的要求,以使它们能够适当地利用,强调了允许fpga在分布式系统中作为成熟的对等体而不是附加到CPU上的必要性。我们认为这个模型需要一个可靠的、无连接的、硬件卸载的传输来支持全局内存空间。我们的结果表明,与基于软件的传输相比,我们的完全成熟的硬件实现如何将延迟改进高达25%,并证明我们的解决方案可以在HPC工作负载(如矩阵-矩阵乘法)中优于目前的状态,从而实现10%的高计算吞吐量。
HPC architects are currently facing myriad challenges from ever tighter power constraints and changing workload characteristics. In this article, we discuss the current state of FPGAs within HPC systems. Recent technological advances show that they are well placed for penetration into the HPC market. However, there are still a number of research problems to overcome; we address the requirements for system architectures and interconnects to enable their proper exploitation, highlighting the necessity of allowing FPGAs to act as full-fledged peers within a distributed system rather than attached to the CPU. We argue that this model requires a reliable, connectionless, hardware-offloaded transport supporting a global memory space. Our results show how our fully fledged hardware implementation gives latency improvements of up to 25% versus a software-based transport, and demonstrates that our solution can outperform the state of the art in HPC workloads such as matrix-matrix multiplication achieving a 10% higher computing throughput.