FReaC Cache: Folded-logic Reconfigurable Computing in the Last Level Cache

FReaC Cache: Folded-logic Reconfigurable Computing in the Last Level Cache
复制标题

DOI:
10.1109/micro50266.2020.00021
复制
发表时间:
2020-10
期刊:
2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Ashutosh Dhar;Xiaohao Wang;H. Franke;Jinjun Xiong;Jian Huang;Wen-mei W. Hwu;N. Kim;Deming Chen
Ashutosh Dhar;Xiaohao Wang;H. Franke;Jinjun Xiong;Jian Huang;Wen-mei W. Hwu;N. Kim;Deming Chen
中科院分区:
其他
文献类型:
--
作者:
Ashutosh Dhar;Xiaohao Wang;H. Franke;Jinjun Xiong;Jian Huang;Wen-mei W. Hwu;N. Kim;Deming Chen

文献摘要

被引文献

相似文献

对更高能效的需求导致了跨平台加速器的激增,边缘设备和云服务器都采用了定制和可重构加速器。然而,现有的解决方案在为加速器提供对工作集的低延迟、高带宽访问方面存在不足,并且存在数据传输的高延迟和能耗问题。这样的成本会严重限制可以加速的任务的最小粒度,从而限制加速器的适用性。在这项工作中,我们提出了FReaC缓存,这是一种新颖的架构,它在最后一级缓存(LLC)中原生支持可重构计算,从而为节能加速器提供低延迟、高带宽的工作集访问。通过利用缓存现有的密集内存阵列、总线和逻辑折叠,我们在LLC中构建了一个可重构的结构,对系统、处理器、缓存和内存体系结构进行了最小的更改。FReaC高速缓存是一种低延迟、低成本、低功耗的芯片外加速器替代品,也是一种灵活、低成本的固定功能加速器替代品。我们展示了与边缘级多核CPU相比,它的平均速度提高了3倍,Perf/W提高了6.1倍,每个缓存片的面积开销增加了3.5%到15.3%。
The need for higher energy efficiency has resulted in the proliferation of accelerators across platforms, with custom and reconfigurable accelerators adopted in both edge devices and cloud servers. However, existing solutions fall short in providing accelerators with low-latency, high-bandwidth access to the working set and suffer from the high latency and energy cost of data transfers. Such costs can severely limit the smallest granularity of the tasks that can be accelerated and thus the applicability of the accelerators. In this work, we present FReaC Cache, a novel architecture that natively supports reconfigurable computing in the last level cache (LLC), thereby giving energy-efficient accelerators low-latency, high-bandwidth access to the working set. By leveraging the cache’s existing dense memory arrays, buses, and logic folding, we construct a reconfigurable fabric in the LLC with minimal changes to the system, processor, cache, and memory architecture. FReaC Cache is a low-latency, low-cost, and low-power alternative to off-die/offchip accelerators, and a flexible, and low-cost alternative to fixed function accelerators. We demonstrate an average speedup of 3X and Perf/W improvements of 6.1X over an edge-class multi-core CPU, and add 3.5% to 15.3% area overhead per cache slice.