CHIPS: Custom Hardware Instruction Processor Synthesis

CHIPS: Custom Hardware Instruction Processor Synthesis
复制标题

DOI:
10.1109/tcad.2008.915536
复制
发表时间:
2008-03
影响因子:
2.9
通讯作者:
K. Atasu;Can C. Özturan;Günhan Dündar;O. Mencer;W. Luk
K. Atasu;Can C. Özturan;Günhan Dündar;O. Mencer;W. Luk
中科院分区:
计算机科学3区
文献类型:
--
作者:
K. Atasu;Can C. Özturan;Günhan Dündar;O. Mencer;W. Luk

文献摘要

被引文献

相似文献

本文描述了一种基于整数线性规划(ILP)的系统,称为定制硬件指令处理器合成(CHIPS),它在给定可用数据带宽和定制逻辑与具有体系结构可见状态寄存器的基线处理器之间的传输延迟的情况下,识别关键代码段的定制指令。我们的方法使设计者能够有选择地限制定制指令的输入和输出操作数的数量。我们描述了一个设计流程,以确定有希望的面积、性能和代码大小之间的权衡。我们研究了输入/输出约束、寄存器堆端口和编译器转换(如IF转换)的影响。我们的实验表明,在大多数情况下,当输入/输出约束被去除时,具有最高性能的解被识别出来。然而,输入/输出约束有助于我们的算法识别频繁使用的代码段,从而减少总体面积开销。给出了11个涵盖密码学和多媒体的基准测试结果,其加速倍数在1.7到6.6倍之间,代码大小减少在6%到72%之间,面积成本在12到256个加法器之间进行最大加速。我们基于ILP的方法具有很好的伸缩性:基本块包含1000多条指令的基准测试可以以最佳方式解决,大多数情况下只需几秒钟。
This paper describes an integer-linear-programming (ILP)-based system called custom hardware instruction processor synthesis (CHIPS) that identifies custom instructions for critical code segments, given the available data bandwidth and transfer latencies between custom logic and a baseline processor with architecturally visible state registers. Our approach enables designers to optionally constrain the number of input and output operands for custom instructions. We describe a design flow to identify promising area, performance, and code-size tradeoffs. We study the effect of input/output constraints, register-file ports, and compiler transformations such as if-conversion. Our experiments show that, in most cases, the solutions with the highest performance are identified when the input/output constraints are removed. However, input/output constraints help our algorithms identify frequently used code segments, reducing the overall area overhead. Results for 11 benchmarks covering cryptography and multimedia are shown, with speed-ups between 1.7 and 6.6 times, code-size reductions between 6% and 72%, and area costs ranging between 12 and 256 adders for maximum speed-up. Our ILP-based approach scales well: benchmarks with basic blocks consisting of more than 1000 instructions can be optimally solved, most of the time within a few seconds.