Hardware Optimizations of Dense Binary Hyperdimensional Computing: Rematerialization of Hypervectors, Binarized Bundling, and Combinational Associative Memory

Hardware Optimizations of Dense Binary Hyperdimensional Computing: Rematerialization of Hypervectors, Binarized Bundling, and Combinational Associative Memory
复制标题

DOI:
10.1145/3314326
复制
发表时间:
2018-07
期刊:
ACM Journal on Emerging Technologies in Computing Systems (JETC)
影响因子:
--
通讯作者:
Manuel Schmuck;L. Benini;Abbas Rahimi
Manuel Schmuck;L. Benini;Abbas Rahimi
中科院分区:
其他
文献类型:
--
作者:
Manuel Schmuck;L. Benini;Abbas Rahimi

文献摘要

被引文献

相似文献

大脑启发的超维(HD)计算用超维空间的点来模拟大脑回路非常大的神经活动模式,也就是说,使用超矢量。超矢量是具有独立同分布(I.I.D.)的D维(伪)随机矢量。构成超宽全息字的分量:例如,D=10,000比特。高清计算的核心是操纵一组种子超向量来构建代表感兴趣对象的复合超向量。它需要通过简单的操作进行内存优化,以实现高效的硬件实现。在这篇文章中,我们提出了在可综合的开源VHDL库中优化HD计算的硬件技术,以使学习和分类任务能够在Xilinx UltraScale FPGA的一小部分上协同执行:(1)我们提出了简单的逻辑操作来动态地重新物化超向量,而不是从内存中加载它们。这些操作通过直接计算其各个种子超向量不需要存储在存储器中的组合超向量来极大地减少存储器占用。(2)随着时间的推移捆绑一系列超矢量需要对每个超矢量分量使用多位计数器。相反,我们建议使用二进制化的背靠背捆绑,而不需要任何计数器。这真正实现了用最少的资源进行片上学习,因为每个超矢量分量在训练过程中保持二进制,以避免其他多位分量。(3)对于每个分类事件,联想记忆通过使用距离度量来寻找一组学习的超向量与查询超向量之间的最接近匹配。该算子与超向量维(D)成正比,因此每个分类事件可能需要O(D)个周期。相应地,我们提出了联想记忆,将分类延迟稳定地减少到单个周期的极端,从而显著提高了分类的吞吐量。(4)以可穿戴生物信号处理应用为例,结合所提出的技术在可穿戴生物信号处理应用中进行了设计空间探索。我们的技术实现了高达2.39倍的面积节省,或2,337倍的吞吐量提高。帕累托最优高清架构仅被映射到18,340个可配置逻辑块(CLB)上,以使用四个肌电传感器学习和分类五个手势。
Brain-inspired hyperdimensional (HD) computing models neural activity patterns of the very size of the brain’s circuits with points of a hyperdimensional space, that is, with hypervectors. Hypervectors are D-dimensional (pseudo)random vectors with independent and identically distributed (i.i.d.) components constituting ultra-wide holographic words: D=10,000 bits, for instance. At its very core, HD computing manipulates a set of seed hypervectors to build composite hypervectors representing objects of interest. It demands memory optimizations with simple operations for an efficient hardware realization. In this article, we propose hardware techniques for optimizations of HD computing, in a synthesizable open-source VHDL library, to enable co-located implementation of both learning and classification tasks on only a small portion of Xilinx UltraScale FPGAs: (1) We propose simple logical operations to rematerialize the hypervectors on the fly rather than loading them from memory. These operations massively reduce the memory footprint by directly computing the composite hypervectors whose individual seed hypervectors do not need to be stored in memory. (2) Bundling a series of hypervectors over time requires a multibit counter per every hypervector component. We instead propose a binarized back-to-back bundling without requiring any counters. This truly enables on-chip learning with minimal resources as every hypervector component remains binary over the course of training to avoid otherwise multibit components. (3) For every classification event, an associative memory is in charge of finding the closest match between a set of learned hypervectors and a query hypervector by using a distance metric. This operator is proportional to hypervector dimension (D), and hence may take O(D) cycles per classification event. Accordingly, we significantly improve the throughput of classification by proposing associative memories that steadily reduce the latency of classification to the extreme of a single cycle. (4) We perform a design space exploration incorporating the proposed techniques on FPGAs for a wearable biosignal processing application as a case study. Our techniques achieve up to 2.39× area saving, or 2,337× throughput improvement. The Pareto optimal HD architecture is mapped on only 18,340 configurable logic blocks (CLBs) to learn and classify five hand gestures using four electromyography sensors.