New Code Generation Algorithm for QueueCore An Embedded Processor with High ILP

New Code Generation Algorithm for QueueCore An Embedded Processor with High ILP
复制标题

DOI:
10.1109/pdcat.2007.12
复制
发表时间:
2007-12
期刊:
Eighth International Conference on Parallel and Distributed Computing, Applications and Technologies (PDCAT 2007)
影响因子:
--
通讯作者:
A. Canedo;B. Abderazek;M. Sowa
A. Canedo;B. Abderazek;M. Sowa
中科院分区:
其他
文献类型:
--
作者:
A. Canedo;B. Abderazek;M. Sowa

文献摘要

相似文献

现代体系结构依赖于利用在指令级发现的并行性来实现高性能。积极的ILP编译器暴露了大量的指令级并行性,在某些情况下,架构寄存器的数量不足以保存潜在并行指令的结果。本文提出了一种新的代码生成方案,为32位的基于嵌入式系统的架构,能够执行大量的ILP。内核的指令隐式地读取它们的操作数并写入结果。编译编译器要求所有指令最多有一个显式操作数,表示为编译时计算的偏移量。此外,指令必须按级别顺序进行调度。该算法成功地限制了所有指令最多只能有一个偏移量参考,计算偏移量值,并对程序进行了分级调度。为了评估新的代码生成方案的有效性,我们开发了一个队列编译器,并编译了一组基准程序。我们的研究结果表明,代码具有更多的并行性比优化的RISC代码的因素范围从1.12到2.30。CoreCore的指令集允许我们生成比优化的RISC代码密度约40%-18%的代码。
Modern architectures rely on exploiting parallelism found at the instruction level to achieve high performance. Aggressive ILP compilers expose high amounts of instruction level parallelism where, in some cases, the number of architected registers is not enough to hold the results of potential parallel instructions. This paper presents a new code generation scheme for the QueueCore, a 32-bit queue-based architecture capable of executing high amounts of ILP. QueueCore's instructions implicitly read their operands and write results. Compiling for the QueueCore requires that all instructions have at most one explicit operand represented as an offset calculated at compile-time. Additionally, the instructions must be scheduled in level-order manner. The proposed algorithm successfully restricts all instructions to have at most one offset reference, it computes the offset values, and makes a level-order scheduling of the program. To evaluate the effectiveness of the new code generation scheme we developed a queue compiler and compiled a set of benchmark programs. Our results show that the code has more parallelism than optimized RISC code by factors ranging from 1.12 to 2.30. QueueCore's instruction set allows us to generate code about 40%-18% denser than optimized RISC code.