Fast and Cycle-Accurate Emulation of Large-Scale Networks-on-Chip Using a Single FPGA

Fast and Cycle-Accurate Emulation of Large-Scale Networks-on-Chip Using a Single FPGA
复制标题

使用单个 FPGA 对大规模片上网络进行快速且周期精确的仿真

DOI:
10.1145/3151758
复制
发表时间:
2017
影响因子:
2.3
通讯作者:
Shimpei Sato and Kenji Kise
Shimpei Sato and Kenji Kise
中科院分区:
计算机科学3区
文献类型:
--
作者:
Thiem Van Chu;Shimpei Sato and Kenji Kise

文献摘要

参考文献

被引文献

相似文献

建模和仿真/仿真在新型片上网络 (NoC) 的研究和开发中发挥着重要作用。然而,传统的软件模拟器速度太慢,以至于研究具有数百到数千个核心的新兴多核系统的 NoC 具有挑战性。最先进的基于 FPGA 的 NoC 仿真器在加速 NoC 仿真方面显示出了巨大的潜力,但由于 FPGA 容量的限制,它们无法仿真大规模的 NoC。此外,在 FPGA 上的综合工作负载下仿真大规模 NoC 通常需要大量内存,因此需要使用片外存储器,这使得整体设计更加复杂,并可能大幅降低仿真速度。本文介绍了使用单个 FPGA 快速、周期精确地仿真具有多达数千个节点的 NoC 的方法。我们首先描述如何仅使用 FPGA 片上存储器 (BRAM) 在合成工作负载下模拟 NoC。接下来,我们提出时分复用的一种新颖用途,其中 BRAM 可有效地使用少量节点来模拟网络,从而克服 FPGA 容量限制。我们提出了模拟直接和间接网络的方法,重点关注常用的网格和胖树(k-aryn-trees)。这与之前仅考虑直接网络的工作不同。使用所提出的方法,我们构建了一个名为 FNoC 的 NoC 仿真器,并演示了对一些具有规范路由器架构的基于网格和基于胖树的 NoC 的仿真。我们的评估结果表明:(1)可仿真的最大NoC的大小仅取决于FPGA片上存储器容量; (2) 可以使用单个 Virtex-7 FPGA 模拟具有 16,384 个节点的基于网状的 NoC(128×128 NoC)以及具有 6,144 个交换节点和 4,096 个终端节点的基于胖树的 NoC(4 进制 6 树 NoC); (3) 在仿真这​​两个 NoC 时,我们分别比 BookSim(最广泛使用的基于软件的 NoC 模拟器之一)实现了 5,047 倍和 232 倍的加速,同时保持了相同的精度水平。
Modeling and simulation/emulation play a major role in research and development of novel Networks-on-Chip (NoCs). However, conventional software simulators are so slow that studying NoCs for emerging many-core systems with hundreds to thousands of cores is challenging. State-of-the-art FPGA-based NoC emulators have shown great potential in speeding up the NoC simulation, but they cannot emulate large-scale NoCs due to the FPGA capacity constraints. Moreover, emulating large-scale NoCs under synthetic workloads on FPGAs typically requires a large amount of memory and thus involves the use of off-chip memory, which makes the overall design much more complicated and may substantially degrade the emulation speed. This article presents methods for fast and cycle-accurate emulation of NoCs with up to thousands of nodes using a single FPGA. We first describe how to emulate a NoC under a synthetic workload using only FPGA on-chip memory (BRAMs). We next present a novel use of time-division multiplexing where BRAMs are effectively used for emulating a network using a small number of nodes, thereby overcoming the FPGA capacity constraints. We propose methods for emulating both direct and indirect networks, focusing on the commonly used meshes and fat-trees (k-aryn-trees). This is different from prior work that considers only direct networks. Using the proposed methods, we build a NoC emulator, called FNoC, and demonstrate the emulation of some mesh-based and fat-tree-based NoCs with canonical router architectures. Our evaluation results show that (1) the size of the largest NoC that can be emulated depends on only the FPGA on-chip memory capacity; (2) a mesh-based NoC with 16,384 nodes (128×128 NoC) and a fat-tree-based NoC with 6,144 switch nodes and 4,096 terminal nodes (4-ary 6-tree NoC) can be emulated using a single Virtex-7 FPGA; and (3) when emulating these two NoCs, we achieve, respectively, 5,047× and 232× speedups over BookSim, one of the most widely used software-based NoC simulators, while maintaining the same level of accuracy.
基于快速可扩展 FPGA 的片上网络仿真模型
DOI: 10.1109/memcod.2011.5970513
发表时间: 2011
期刊: Ninth ACM/IEEE International Conference on Formal Methods and Models for Codesign (MEMPCODE2011)
影响因子: --
作者:
Michael Papamichael
通讯作者: Michael Papamichael
FPGA 的技术扩展:应用和架构的趋势
DOI: 10.1109/fccm.2015.11
发表时间: 2015
期刊: 2015 IEEE 23rd Annual International Symposium on Field-Programmable Custom Computing Machines
影响因子: --
作者:
Lesley Shannon;V. Cojocaru;Cong Nguyen Dao;P. Leong
通讯作者: P. Leong
多核处理器的快速且周期精确的建模
DOI: 10.1109/ispass.2012.6189224
发表时间: 2012
期刊: 2012 IEEE International Symposium on Performance Analysis of Systems & Software
影响因子: --
作者:
Asif Khan;M. Vijayaraghavan;Silas Boyd;Arvind
通讯作者: Arvind
FPGA 上多核处理器的周期精确建模
DOI: --
发表时间: 2013
期刊: --
影响因子: --
作者:
Asif Khan
通讯作者: Asif Khan