Main memory latency simulation: the missing link

Main memory latency simulation: the missing link
复制标题

主内存延迟模拟:缺失的环节

DOI:
10.1145/3240302.3240317
复制
发表时间:
2018
期刊:
MEMSYS 2018
影响因子:
--
通讯作者:
Jacob, Bruce
Jacob, Bruce
中科院分区:
--
文献类型:
--
作者:
Verdejo, Rommel Sánchez;Asifuzzaman, Kazi;Radulovic, Milan;Radojković, Petar;Ayguadé, Eduard;Jacob, Bruce

文献摘要

参考文献

被引文献

相似文献

社区接受了对主存进行详细模拟的需要。目前,CPU模拟器通常与周期精确的主存模拟器耦合。然而,耦合CPU和内存模拟器并不是一项简单的任务,因为末级缓存和内存DIMM之间的某些电路很容易被忽视,因此无法解释。在本文中,我们采取了一种方法来量化丢失的周期在主内存模拟中。为此,我们执行了内存密集型微基准测试,以验证基于Intel桑迪Bridge E5-2670系统的DRAM和DRAMsim 2建模的仿真基础设施。我们在真实的桑迪桥E5-2670机器上执行相同的微基准测试,识别模拟器测量中缺失的20 ns。这是一个巨大的差异,在所研究的系统中,对应于总主存延迟的三分之一。我们提出了多种方案来在仿真模型中添加额外的延迟,以解决缺失的周期。此外,我们使用SPEC CPU 2006基准验证的建议。最后,我们在七个主流和新兴的计算平台上重复了主存延迟测量。我们的结果表明,末级缓存(LLC)和主内存之间的延迟范围在数十到数百纳秒之间,因此我们强调在执行任何测量之前在系统模拟器中适当调整和验证这些参数。总的来说,我们相信这项研究将改善主存模拟,从而更好地整体系统分析和探索在计算机体系结构社区进行。
The community accepted the need for a detailed simulation of main memory. Currently, the CPU simulators are usually coupled with the cycle-accurate main memory simulators. However, coupling CPU and memory simulators is not a straight-forward task because some pieces of the circuitry between the last level cache and the memory DIMMs could be easily overlooked and therefore not accounted for.In this paper, we take an approach to quantify the missing cycles in the main memory simulation. To that end, we execute a memory intensive microbenchmark to validate a simulation infrastructure based on ZSim and DRAMsim2 modeling an Intel Sandy Bridge E5-2670 system. We execute the same microbenchmark on a real Sandy Bridge E5-2670 machine identifying a missing 20 ns in the simulator measurements. This is a huge difference that, in the system under study, corresponds to one-third of the overall main memory latency. We propose multiple schemes to add an extra delay in the simulation model to account for the missing cycles. Furthermore, we validate the proposals using the SPEC CPU2006 benchmarks. Finally, we repeat the main memory latency measurements on seven mainstream and emerging computing platforms. Our results show that latency between the Last Level Cache (LLC) and the main memory ranges between tens and hundreds of nanoseconds, so we emphasize on properly adjust and validate these parameters in system simulators before any measurements are performed. Overall, we believe this study would improve main memory simulation leading to the better overall system analysis and explorations performed in the computer architecture community.
用于硬件模拟器详细验证和调整的微基准测试
DOI: --
发表时间: 2017
期刊: International Symposium on High Performance Computing Systems and Applications
影响因子: --
作者:
Rommel Sánchez Verdejo;Petar Radojković
通讯作者: Petar Radojković