Trireme: Exploration of Hierarchical Multi-level Parallelism for Hardware Acceleration

Trireme: Exploration of Hierarchical Multi-level Parallelism for Hardware Acceleration
复制标题

DOI:
10.1145/3580394
复制
发表时间:
2023-01
影响因子:
2
通讯作者:
Georgios Zacharopoulos;Adel Ejjeh;Ying Jing;En-Yu Yang;Tianyu Jia;I. Brumar;Jeremy Intan;Muhammad Huzaifa;S. Adve;Vikram S. Adve;Gu-Yeon Wei;D. Brooks
Georgios Zacharopoulos;Adel Ejjeh;Ying Jing;En-Yu Yang;Tianyu Jia;I. Brumar;Jeremy Intan;Muhammad Huzaifa;S. Adve;Vikram S. Adve;Gu-Yeon Wei;D. Brooks
中科院分区:
计算机科学3区
文献类型:
--
作者:
Georgios Zacharopoulos;Adel Ejjeh;Ying Jing;En-Yu Yang;Tianyu Jia;I. Brumar;Jeremy Intan;Muhammad Huzaifa;S. Adve;Vikram S. Adve;Gu-Yeon Wei;D. Brooks

文献摘要

相似文献

包括特定域的加速器的异质系统的设计是一个挑战和耗时的过程。例如扩展现实(XR)为各种形式的并行执行提供了机会,包括循环级别,任务级别和管道并行性,以协助设计过程并揭示所有可能的并行性,我们提出了Trireme,这是一个完全自动化的工具链FPGA SOC被用作目标平台,使用XR域中的基准,弹射式HLS [7]用于合成RTL。与仅软件实现相比,较小的应用程序最多可达37倍。
The design of heterogeneous systems that include domain specific accelerators is a challenging and time-consuming process. While taking into account area constraints, designers must decide which parts of an application to accelerate in hardware and which to leave in software. Moreover, applications in domains such as Extended Reality (XR) offer opportunities for various forms of parallel execution, including loop level, task level, and pipeline parallelism. To assist the design process and expose every possible level of parallelism, we present Trireme, a fully automated tool-chain that explores multiple levels of parallelism and produces domain-specific accelerator designs and configurations that maximize performance, given an area budget. FPGA SoCs were used as target platforms, and Catapult HLS [7] was used to synthesize RTL using a commercial 12 nm FinFET technology. Experiments on demanding benchmarks from the XR domain revealed a speedup of up to 20×, as well as a speedup of up to 37× for smaller applications, compared to software-only implementations.