Snafu: An Ultra-Low-Power, Energy-Minimal CGRA-Generation Framework and Architecture

Snafu: An Ultra-Low-Power, Energy-Minimal CGRA-Generation Framework and Architecture
复制标题

DOI:
10.1109/isca52012.2021.00084
复制
发表时间:
2021-06
期刊:
2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Graham Gobieski;A. Atli;K. Mai;Brandon Lucia;Nathan Beckmann
Graham Gobieski;A. Atli;K. Mai;Brandon Lucia;Nathan Beckmann
中科院分区:
其他
文献类型:
--
作者:
Graham Gobieski;A. Atli;K. Mai;Brandon Lucia;Nathan Beckmann

文献摘要

相似文献

超低功耗(ULP)器件正变得越来越普遍,使许多新兴的传感应用成为可能。能效在这些应用中至关重要,因为能效决定了电池供电部署中的设备寿命和能量收集部署中的性能。不幸的是,现有的设计不足,因为ASIC的前期成本太高,以前的ULP架构太低效或不灵活。我们提出Snafu,第一个框架,灵活地产生ULP粗粒度可重构阵列(CGRAs)。Snafu为处理元件(PE)提供了标准接口,使其能够轻松集成用于新应用的新型PE。与之前的高性能、高功率CGRA不同,Snafu的设计从根本上降低了能耗,同时最大限度地提高了灵活性。Snafu节省能源配置PE和路由器为一个单一的操作,以最大限度地减少交换活动,通过最大限度地减少缓冲内的结构,通过实现静态路由,无缓冲,多跳网络,并通过执行操作,以避免昂贵的tag-token matching.We进一步提出Snafu-Arch,一个完整的ULP系统,集成了一个实例化的Snafu结构旁边的标量RISC-V核心和内存。我们在RTL中实现Snafu,并在一套常见的传感基准测试中在工业亚28 nm FinFET工艺上对其进行评估。Snafu-Arch的工作功率<1 mW,比大多数现有CGRA低几个数量级。Snafu-Arch比现有的最先进的通用ULP架构节省41%的能源,运行速度快4.4倍。此外,我们进行了三个全面的案例研究,以量化Snafu的可编程性的成本。我们发现Snafu-Arch接近于采用相同技术构建的ASIC设计,平均仅使用2.6倍的能量。
Ultra-low-power (ULP) devices are becoming pervasive, enabling many emerging sensing applications. Energy-efficiency is paramount in these applications, as efficiency determines device lifetime in battery-powered deployments and performance in energy-harvesting deployments. Unfortunately, existing designs fall short because ASICs’ upfront costs are too high and prior ULP architectures are too inefficient or inflexible.We present Snafu, the first framework to flexibly generate ULP coarse-grain reconfigurable arrays (CGRAs). Snafu provides a standard interface for processing elements (PE), making it easy to integrate new types of PEs for new applications. Unlike prior high-performance, high-power CGRAs, Snafu is designed from the ground up to minimize energy consumption while maximizing flexibility. Snafu saves energy by configuring PEs and routers for a single operation to minimize switching activity; by minimizing buffering within the fabric; by implementing a statically routed, bufferless, multi-hop network; and by executing operations in-order to avoid expensive tag-token matching.We further present Snafu-Arch, a complete ULP system that integrates an instantiation of the Snafu fabric alongside a scalar RISC-V core and memory. We implement Snafu in RTL and evaluate it on an industrial sub-28 nm FinFET process across a suite of common sensing benchmarks. Snafu-Arch operates at <1 mW, orders-of-magnitude less power than most prior CGRAs. Snafu-Arch uses 41% less energy and runs 4.4× faster than the prior state-of-the-art general-purpose ULP architecture. Moreover, we conduct three comprehensive case-studies to quantify the cost of programmability in Snafu. We find that Snafu-Arch is close to ASIC designs built in the same technology, using just 2.6× more energy on average.