Dynamic Instructions Issue Algorithm and a Queue Execution Model toward the Design of Hybrid Processor Architecture

Dynamic Instructions Issue Algorithm and a Queue Execution Model toward the Design of Hybrid Processor Architecture
复制标题

混合处理器架构设计的动态指令发布算法和队列执行模型

DOI:
--
复制
发表时间:
2014
期刊:
--
影响因子:
--
通讯作者:
Abderazek Ben Abdallah
Abderazek Ben Abdallah
中科院分区:
--
文献类型:
--
作者:
Abderazek Ben Abdallah

文献摘要

被引文献

相似文献

微处理器体系结构的演变取决于不断变化的技术方面。随着芯片密度和速度的提高,支持的指令集、兼容性和硬件创新在定义架构权衡中变得越来越重要。为了满足日益增长的对更高级别计算能力的需求,计算机架构师需要研究在考虑不断变化的技术和应用的同时继续提高微处理器性能的技术。传统处理器可以通过从原始指令流发出无序(O00)指令来实现提高的性能。实现OOO指令发布方案需要一种强大的机制来防止错误执行的指令更新寄存器值并提高性能。此外,如果指令之间出现依赖关系、分支或陷阱,则性能会降低。在本文中,我们首先提出了一种新的算法,称为动态快速发布(DFI)机制,以无序(OOO)方式向多个并行功能单元发布指令。该系统解决了数据依赖问题,支持精确的中断和分支预测,这是超标量机器中指令动态调度的主要问题。结果只写入一次,直接写入寄存器文件(RF)。为了确保结果按顺序写入其相应的输出寄存器中,指令顺序和状态的记录由状态缓冲器(STB)维护。为了从中断或分支未命中预测中恢复处理器状态,实现了状态缓冲器(STB)和恢复列表表格(RLT)。我们的性能评估结果表明,对于具有不同IPC级别和较高分支错误率的程序,DFI比传统的Super-Base结构获得了12%的平均增益。其次,我们提出了一种直接结合硬件队列作为其主要数据处理结构的执行模型体系结构。我们提出了一个QEM指令生成的数学模型。然后,我们证明了队列执行模型的指令序列可以正确和容易地从解析树生成,并且可以用于计算任何任意表达式。除了数学模型之外,我们还给出了上述执行模型的新颖方面以及架构背后的原理。为了试验我们的体系结构创新并评估其效率,我们在软件中构建体系结构。设计了一个模拟器,用于确定一些基准程序的循环计数性能。上述队列执行模型和DFI系统旨在在一个所谓的功能分配寄存器微处理器中实现,该微处理器包含多编程环境,结合了QEM和RISC计算模型的最佳未来。最后给出了整个处理器的新颖之处以及该体系结构的基本原理。
The evolution of microprocessor architecture depends upon the changing aspect of technology. As die density and speed increase, supported instructions set, compatibility and hardware innovations become increasingly important in defining architecture trade-offs. To satisfy the ever growing need for higher levels of computing power, computer architects need, then, to investigate techniques that continue improving the performance of microprocessors while considering both changing technology and applications. Conventional processors can achieve increased performance by issuing instructions Out-of-Order (OoO) from the original instruction stream. Implementing an OoO instruction issue scheme requires a powerful mechanism to prevent incorrectly executed instructions from updating registers values and to boost performance. In addition, performance degrades if dependencies, branches or traps among instructions appear. In this thesis, we first propose a new algorithm named Dynamic Fast Issue (DFI) mechanism to issue instructions in an out-of-order (OoO) scheme to multiple parallel functional units. The above system solves data dependencies, supports precise interrupt and branch prediction, which are the main problems associated with the dynamic scheduling of instructions in superscalar machines. Results are written only once, directly into the register file (RF). To ensure that results are written in order in their appropriate output registers, a record of instruction order and state is maintained by a status buffer (STB). To recover the processor state from an interrupt or a branch miss-prediction, a status buffer (STB) and a recovery list table (RLT) are implemented. Our performances evaluation results show that the DFI achieves a 12% average gain over the SUPER-Base conventional architecture for programs with varying level of IPC and high branch miss-predictions rates. Second, we propose an execution model's architecture that directly incorporates hardware Queue as its primarily data handling structure. We present a mathematical model for the QEM's instructions generation. We prove, then, that the Queue execution model’s instructions sequence can be correctly and easily generated from a parse tree and can be used to evaluate any arbitrary expression. In addition to the mathematical model, we give the novel aspects of the above execution model as well as the principle underlying the architecture. To experiment with our architecture innovation and to evaluate its efficiency, we build the architecture in software. A simulator, which determines the cycle count performance for some benchmark programs, was designed. The above Queue execution model and the DFI system are intended to be implemented in a so called Functional Assignment Register Microprocessor that embraces multi programming environments, combining the best futures of QEM and RISC models of computing. The novel aspects of the whole processor as well as the principle underlying the architecture are finally given.