EVE: Ephemeral Vector Engines

EVE: Ephemeral Vector Engines
复制标题

DOI:
10.1109/hpca56546.2023.10071074
复制
发表时间:
2023-02
期刊:
2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Khalid Al-Hawaj;T. Ta;Nick Cebry;Shady O. Agwa;O. Afuye;Eric Hall;Courtney Golden;A. Apsel;C. Batten
Khalid Al-Hawaj;T. Ta;Nick Cebry;Shady O. Agwa;O. Afuye;Eric Hall;Courtney Golden;A. Apsel;C. Batten
中科院分区:
其他
文献类型:
--
作者:
Khalid Al-Hawaj;T. Ta;Nick Cebry;Shady O. Agwa;O. Afuye;Eric Hall;Courtney Golden;A. Apsel;C. Batten

文献摘要

被引文献

相似文献

最近在主流指令集架构中采用向量扩展,这证明了对向量架构的兴趣重新抬头。传统上,向量引擎通过利用其固有的规律性来利用这种抽象来提高性能和效率。基于SRAM的内存计算的最新工作已经显示出减少这些引擎的面积开销的希望。在这项工作中,我们提出了短暂的向量引擎(EVE),我们利用基于SRAM的内存计算技术以及位外设计算,以促进高效的向量执行。EVE使用了一种新颖的位混合执行方法,在吞吐量和延迟之间取得了平衡。在Rodinia和RiVEC基准套件上进行评估,与乱序处理器相比,EVE实现了近8倍的速度提升,与集成矢量单元相比,EVE实现了4.59倍的速度提升。EVE实现了与主动解耦矢量单元相当的速度提升,并将面积归一化性能提高了2倍以上。通过重新利用L2缓存中的SRAM阵列来创建短暂的向量执行单元,EVE能够有效地实现高性能,同时仅产生11.7%的面积开销。
There has been a resurgence of interest in vector architectures evident by recent adoption of vector extensions in mainstream instruction set architectures. Traditionally, vector engines leverage this abstraction by exploiting its inherent regularity to increase performance and efficiency. Recent work on SRAM-based compute-in-memory has shown promise in reducing the area overhead of these engines. In this work, we propose ephemeral vector engines (EVE) where we leverage SRAM-based compute-in-memory techniquesas well as bit-peripheral computations to facilitate efficient vector execution. EVE uses a novel approach of bit-hybrid execution, striking a balance between throughput and latency. Evaluated on the Rodinia and RiVEC benchmark suites, EVE achieves almost 8× speed-up compared to an out-of-order processor and 4.59× compared to an integrated vector unit. EVE achieves speed-ups comparable to an aggressive decoupled vector unit and increases the area-normalized performance by over 2 ×. By repurposing SRAM arrays in the L2 cache to create ephemeral vector execution units, EVE is able to efficiently achieve high performance while incurring as little as 11.7% area overhead.