The multicluster architecture: reducing cycle time through partitioning

The multicluster architecture: reducing cycle time through partitioning
复制标题

多集群架构:通过分区减少周期时间

DOI:
10.1109/micro.1997.645806
复制
发表时间:
1997
期刊:
Proceedings of 30th Annual International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Z. Vranesic
Z. Vranesic
中科院分区:
--
文献类型:
--
作者:
K. Farkas;P. Chow;N. Jouppi;Z. Vranesic

文献摘要

被引文献

相似文献

我们介绍的多集群体系结构提供了一种分散的、动态调度的体系结构,其中该体系结构的寄存器堆、调度队列和功能单元分布在多个集群中,并且每个集群都被分配了体系结构寄存器的一个子集。与具有相同数量硬件资源的单群集体系结构相比,多群集体系结构的动机是通过减少关键定时路径上组件的大小和复杂性来减少时钟周期时间。然而,资源分区会引入指令执行开销,并可能减少并发执行的指令数量。针对分区的这两个负面副产品,我们提出了一种静态指令调度算法。我们描述了该算法,并使用跟踪驱动的SPEC92基准程序仿真来评估其有效性。该评估表明,对于所考虑的配置,多集群体系结构在特征大小低于0.35/SPL MU/m时可能具有显著的性能优势,值得进一步研究。
The multicluster architecture that we introduce offers a decentralized, dynamically scheduled architecture, in which the register files, dispatch queue, and functional units of the architecture are distributed across multiple clusters, and each cluster is assigned a subset of the architectural registers. The motivation for the multicluster architecture is to reduce the clock cycle time, relative to a single-cluster architecture with the same number of hardware resources, by reducing the size and complexity of components on critical timing paths. Resource partitioning, however, introduces instruction-execution overhead and may reduce the number of concurrently executing instructions. To counter these two negative by-products of partitioning, we developed a static instruction scheduling algorithm. We describe this algorithm, and using trace-driven simulations of SPEC92 benchmarks, evaluate its effectiveness. This evaluation indicates that for the configurations considered the multicluster architecture may have significant performance advantages at feature sizes below 0.35 /spl mu/m, and warrants further investigation.