A free-extendible and ultralow-power nonvolatile multi-core associative coprocessor based on MRAM with inter-core pipeline scheme for large-scale full-adaptive nearest pattern searching

A free-extendible and ultralow-power nonvolatile multi-core associative coprocessor based on MRAM with inter-core pipeline scheme for large-scale full-adaptive nearest pattern searching
复制标题

DOI:
10.35848/1347-4065/ab72d0
复制
发表时间:
2020-03
影响因子:
1.5
通讯作者:
Yitao Ma;S. Miura;H. Honjo;S. Ikeda;T. Endoh
Yitao Ma;S. Miura;H. Honjo;S. Ikeda;T. Endoh
中科院分区:
物理与天体物理4区
文献类型:
--
作者:
Yitao Ma;S. Miura;H. Honjo;S. Ikeda;T. Endoh

文献摘要

相似文献

采用开放式设计,设计了一种基于MRAM的非易失性多核相联协处理器,在保持高电路密度、全自适应和高速的同时,实现了较高的功率效率。该协处理器适用于从大规模的正常/预聚类数据集中搜索最近的模式。实现了一种核间流水线操作方案,该方案吸收了核内时间域最小搜索的延迟,保证了协处理器的高速运行,并通过增加核数灵活地扩展了协处理器。还采用了一种自优化的功率门控方案,使运算功率最小化,并在簇数(K)变大时进一步降低运算功率。采用90 nm CMOS/70 nm垂直MTJ混合工艺,设计并制作了12核芯片原型,测试结果表明芯片工作在100 MHz。平均工作功率仅为68μW@K=24,与最新的常规研究相比,功率效率提高了40多倍。
A nonvolatile multi-core associative coprocessor based on the MRAM is developed with open-end design, which achieves the higher power efficiency maintaining the high circuit density, full adaptivity and high-speed at the same time. This proposed coprocessor is proposed applicable for searching nearest pattern from the large-scale normal/pre-clustered data sets. An inter-core pipeline operation scheme is implemented, which absorbs the delay of in-core time-domain minimum searching to ensure the high-speed and flexibly extends the coprocessor by increasing the core count. A self-optimized power gating scheme is also employed, which minimizes the operation power and further reduces it when the cluster number (K) becomes larger. The prototype chip with 12-core is designed and fabricated under 90 nm CMOS/70 nm perpendicular-MTJ hybrid technology, and the chip operation at 100 MHz is demonstrated by measurement results. The average operation power is only 68 μW @K = 24, and more than 40-time higher power efficiency is achieved comparing to the latest conventional researches.