Rapid Generation of High-Quality RISC-V Processors from Functional Instruction Set Specifications

Rapid Generation of High-Quality RISC-V Processors from Functional Instruction Set Specifications
复制标题

根据功能指令集规范快速生成高质量 RISC-V 处理器

DOI:
10.1145/3316781.3317890
复制
发表时间:
2019
期刊:
2019 56th ACM/IEEE Design Automation Conference (DAC)
影响因子:
--
通讯作者:
Zhiru Zhang
Zhiru Zhang
中科院分区:
--
文献类型:
--
作者:
Gai Liu;Joseph Primmer;Zhiru Zhang

文献摘要

参考文献

被引文献

相似文献

诸如人工智能和计算机视觉等新兴领域的计算加速度的普及越来越普及,导致对特定领域的加速器的需求日益增长,这通常是作为执行一组域优化说明的专业处理器实现的。快速探索(1)定制指令集的各种可能性以及(2)其相应的微构造特征对于实现最佳的收集质量(QOR)至关重要。但是,在寄存器传输级别(RTL)处的手动设计过程通常会阻碍这种能力。这种基于RTL的方法通常昂贵且反应缓慢,当时设计规格在指令集级别和/或微构造级别上发生变化。我们在辅助方面解决了特定领域的处理器设计中的这种不足,行为级合成RISC-V处理器的框架。从不合时宜的功能指令集说明中,Assiss生成了实施不同微构造设计选择的RISC-V处理器,这可以在不同的QOR指标之间进行有效的权衡。我们演示了60多个超过60多种订购的处理器实现,其中RISC-V 32i指令集的管道结构不同,其中一些主导了该区域绩效帕累托边境中手动优化的对应物。此外,我们提出了一种基于自动调用的方法,以优化给定性能限制和技术目标下的实现。我们进一步介绍了为密码和机器学习应用程序综合各种自定义说明扩展程序和自定义说明集的案例研究。
The increasing popularity of compute acceleration for emerging domains such as artificial intelligence and computer vision has led to the growing need for domain-specific accelerators, often implemented as specialized processors that execute a set of domain-optimized instructions. The ability to rapidly explore (1) various possibilities of the customized instruction set, and (2) its corresponding micro-architectural features is critical to achieve the best quality-of-results (QoRs). However, this ability is frequently hindered by the manual design process at the register transfer level (RTL). Such an RTL-based methodology is often expensive and slow to react when the design specifications change at the instruction-set level and/or micro-architectural level.We address this deficiency in domain-specific processor design with ASSIST, a behavior-level synthesis framework for RISC-V processors. From an untimed functional instruction set description, ASSIST generates a spectrum of RISC-V processors implementing varying micro-architectural design choices, which enables effective tradeoffs between different QoR metrics. We demonstrate the automatic synthesis of more than 60 in-order processor implementations with varying pipeline structures from the RISC-V 32I instruction set, some of which dominate the manually optimized counterparts in the area-performance Pareto frontier. In addition, we propose an autotuning-based approach for optimizing the implementations under a given performance constraint and the technology target. We further present case studies of synthesizing various custom instruction extensions and customized instruction sets for cryptography and machine learning applications.
DOI: 10.1109/tc.2013.37
发表时间: 2014
影响因子: 3.7
作者:
Mokhov A
通讯作者: Mokhov A