Fast and cycle-accurate modeling of a multicore processor

Fast and cycle-accurate modeling of a multicore processor
复制标题

多核处理器的快速且周期精确的建模

DOI:
10.1109/ispass.2012.6189224
复制
发表时间:
2012
期刊:
2012 IEEE International Symposium on Performance Analysis of Systems & Software
影响因子:
--
通讯作者:
Arvind
Arvind
中科院分区:
--
文献类型:
--
作者:
Asif Khan;M. Vijayaraghavan;Silas Boyd;Arvind

文献摘要

参考文献

被引文献

相似文献

理想的模拟器允许架构师快速探索设计方案并准确确定其对性能的影响。设计探索要求模拟器易于修改,而准确的性能评估需要详细的模型。不幸的是,详细的建模不仅会影响模拟器修改的容易程度,还会影响模拟器执行的速度,导致保真度被仿真速度所取代。尽管基于fpga的模拟器比软件模拟器具有更高的速度,但牺牲保真度仍然是常见的。在本文中,我们提出了Arete,一个基于fpga的处理器模拟器,它提供了高性能,精度和可修改性。我们从多核架构的周期级规范开始,其中包括现实的顺序核心和共享,连贯内存和片上网络的详细模型。然后,我们描述了如何在fpga上忠实有效地实现该规范。Arete每核提供高达11 MIPS的性能。我们在现成的SMP Linux上运行PARSEC基准套件的一个子集,并在8核模型上实现了55 MIPS的平均性能。我们还描述了两个重要的架构探索:一个涉及三个不同的分支预测器,另一个需要对缓存一致性协议进行重大修改。
An ideal simulator allows an architect to swiftly explore design alternatives and accurately determine their impact on performance. Design exploration requires simulators to be easily modifiable, and accurate performance estimates require detailed models. Unfortunately, detailed modeling not only impacts the ease with which a simulator can be modified, but also the speed at which it can be executed, resulting in fidelity being traded for simulation speed. Although FPGA-based simulators have dramatically higher speed than software simulators, sacrificing fidelity is still common. In this paper we present Arete, an FPGA-based processor simulator, which offers high performance along with accuracy and modifiability. We begin with a cycle-level specification of a multicore architecture which includes realistic in-order cores and detailed models of shared, coherent memory and on-chip network. We then describe how this specification is implemented faithfully and efficiently on FPGAs. Arete delivers a performance of up to 11 MIPS per core. We run a subset of the PARSEC benchmark suite on top of off-the-shelf SMP Linux, and achieve an average performance of 55 MIPS for an 8-core model.We also describe two significant architectural explorations: one involving three different branch predictors and the other requiring major modifications to the cache-coherence protocol.
DOI: --
发表时间: 2004
期刊: 第22回日本ロボット学会学術講演会予稿集
影响因子: --
作者:
M.Higashimori;東森 充;東森 充
通讯作者: 東森 充