An Effective Architecture for Trace-Driven Emulation of Networks-on-Chip on FPGAs

An Effective Architecture for Trace-Driven Emulation of Networks-on-Chip on FPGAs
复制标题

DOI:
10.1109/fpl.2018.00078
复制
发表时间:
2018-08
期刊:
2018 28th International Conference on Field Programmable Logic and Applications (FPL)
影响因子:
--
通讯作者:
Thiem Van Chu;Kenji Kise
Thiem Van Chu;Kenji Kise
中科院分区:
其他
文献类型:
--
作者:
Thiem Van Chu;Kenji Kise

文献摘要

相似文献

现代多核系统使用片上网络(NoCs)在其核心之间移动数据。随着核心数量的增加,整体性能对NoC性能变得高度敏感。因此,NoC的研究和开发在设计具有数百到数千个核心的未来系统中发挥着关键作用。然而,就系统复杂性而言,当前用于评估NOC的方法是不可扩展的。传统的软件模拟器对于评估中大型国家石油公司来说太慢了。最新的基于现场可编程门阵列的仿真器提供了比软件仿真器有希望的仿真加速比。然而,由于现场可编程门阵列的容量限制,在现场可编程门阵列上模拟具有数百到数千个节点的大规模片上网络是一个具有挑战性的问题。此外,支持跟踪驱动的仿真并不容易,因为跟踪数据必须存储在FPGA外部(通常存储在片外DRAM中)。现有的大多数基于FPGA的NoC仿真器依赖于MicroBlaze等软处理器或SoC FPGA上的硬件处理器来从片外存储器加载跟踪数据、生成消息、将消息注入目标NoC、操纵仿真并确保没有时序误差。这种方法使实现变得容易,但极大地降低了仿真速度。提出了一种在现场可编程门阵列上实现轨迹驱动的片上网络仿真的有效结构。我们提出了扩展到大型片上网络的方法,并有效地隐藏了片外存储器访问延迟。我们的评估结果表明:(1)与目前应用最广泛的片上网络模拟器BookSim相比,该方案在模拟片上网络时获得了260倍的加速比;(2)在模拟64x64片上网络时,加速比提高到三个数量级。
Modern many-core systems use Networks-on-Chip (NoCs) to move data around their cores. As the number of cores increases, the overall performance becomes highly sensitive to the NoC performance. Research and development of NoCs thus play a key role in designing future systems with hundreds to thousands of cores. However, current methodologies for evaluating NoCs are not scalable with respect to the system complexity. Conventional software simulators are too slow for evaluating middle-and large-scale NoCs. Recent FPGA-based emulators provide promising emulation speedups over software simulators. However, emulating large-scale NoCs with hundreds to thousands of nodes on FPGAs is a challenging problem because of the FPGA capacity constraints. Moreover, supporting trace-driven emulation is not trivial because trace data must be stored outside of the FPGA (usually in off-chip DRAM). Most of the existing FPGA-based NoC emulators rely on soft processors like Microblaze or hard processors on SoC FPGAs for loading trace data from the off-chip memory, generating messages, injecting the messages to the target NoC, manipulating the emulation, and making sure that there is no timing error. This approach makes the implementation easy but drastically degrades the emulation speed. This paper proposes an effective architecture for trace-driven emulation of NoCs on FPGAs. We present methods to scale to large NoCs and effectively hide the off-chip memory access latency. Our evaluation results show that (1) the proposal achieves a speedup of 260x compared to BookSim, one of the most widely used NoC simulators, when emulating an 8x8 NoC with trace data collected from full-system simulation of the PARSEC benchmark suite; and (2) the speedup is increased to three orders of magnitude when emulating a 64x64 NoC.