4.6 A 144Kb Annealing System Composed of 9× 16Kb Annealing Processor Chips with Scalable Chip-to-Chip Connections for Large-Scale Combinatorial Optimization Problems

4.6 A 144Kb Annealing System Composed of 9× 16Kb Annealing Processor Chips with Scalable Chip-to-Chip Connections for Large-Scale Combinatorial Optimization Problems
复制标题

4.6 由 9 个 16Kb 退火处理器芯片组成的 144Kb 退火系统,具有可扩展的芯片到芯片连接,用于解决大规模组合优化问题

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Solid-State Circuits Conference
影响因子:
--
通讯作者:
M. Yamaoka
M. Yamaoka
中科院分区:
--
文献类型:
--
作者:
Takashi Takemoto;Kasho Yamamoto;C. Yoshimura;Masato Hayashi;Masafumi Tada;Hiroaki Saito;Mayumi Mashimo;M. Yamaoka

文献摘要

被引文献

相似文献

在一种新的计算机体系结构上已经取得了实质性的进展,称为退火处理器(AP)[1-4]。该算法通过提供一种快速求解伊辛模型宏观状态的方法,可以有效地解决NP难的组合优化问题。具体地,基于CMOS工艺的各种类型的AP(CMOS-AP)通过利用基于模拟退火(SA)的快速并行自旋更新来显著提高退火系统的可扩展性和功率效率[2] -[4]。CMOS-AP的进一步发展需要克服两个挑战:通过扩展系数的位宽来提高退火处理的精度,以及实现由具有8路连接的AP芯片组成的多芯片退火系统。在这项工作中,我们开发了一个可扩展的CMOS-AP与两个关键技术:(i)一个触发器(FF)为基础的自旋电路,允许可扩展的位宽通过复制的大都会算法,这是SA与一个固定的温度,和(ii)一个芯片间接口(I/F)与数据压缩方法,利用退火特性,以获得多芯片操作,而不会降低退火速度和精度。CMOS-AP演示了9美元的 imes 16 k $旋转系统,退火速度为233美元 10万美元,计算能耗972美元 比在CPU上运行SG 3低10倍。
Substantial progress has been made on a new computer architecture, known as an annealing processor (AP) [1–4]. The AP can effectively solve NP-hard combinatorial optimization problems by providing a fast method for finding the grand state of an Ising model. In particular, various types of APs based on a CMOS process (CMOS-AP) significantly improve the scalability and power efficiency of the annealing system by utilizing fast parallel spin updates on the basis of simulated annealing (SA) [2] –[4]. Further development of CMOS-APs requires overcoming two challenges: improving the accuracy of the annealing processing by expanding the bitwidth of coefficients and attaining a multi-chip annealing system consisting of AP chips with 8-way connectivity. In this work, we developed a scalable CMOS-AP with two key technologies: (i) A flip-flop (FF)-based spin circuit allowing expandable bitwidth by reproducing the Metropolis algorithm, which is SA with a fixed temperature, and (ii) an inter-chip interface (I/F) with a data-compression method utilizing annealing characteristics to obtain multi-chip operation without degrading annealing speed and accuracy. The CMOS-AP demonstrated multi-chip operation of the $9 imes16k$ spin system with an annealing speed $233 imes$ faster and a calculation energy $972 imes$ lower than running SG3 on a CPU.