Analytical Modeling the Multi-Core Shared Cache Behavior With Considerations of Data-Sharing and Coherence

Analytical Modeling the Multi-Core Shared Cache Behavior With Considerations of Data-Sharing and Coherence
复制标题

考虑数据共享和一致性的多核共享缓存行为分析建模

DOI:
10.1109/access.2021.3053350
复制
发表时间:
2020
期刊:
影响因子:
3.9
通讯作者:
Jiancong Ge
Jiancong Ge
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ming Ling;Xiaoqian Lu;Guangmin Wang;Jiancong Ge

文献摘要

参考文献

相似文献

为了缓解日益严重的“功率墙”和“内存墙”问题,具有多级缓存层次结构的多核架构在现代处理器中被广泛接受。然而,架构的复杂性使得共享缓存的建模极其复杂。在本文中,我们提出了一个数据共享感知的分析模型,用于估计多核场景下下游共享缓存的缺失率。为了避免传统方法所需的耗时的缓存架构的完整模拟,所提出的模型还可以与我们改进的上游缓存分析模型集成,该模型还可以评估具有相似精度的最先进方法的相干缺失,而开销仅为十分之一。我们用PARSEC 2.1基准测试套件中的13个应用程序的gem5仿真结果验证了我们的分析模型。与gem5在8种硬件配置(包括双核和四核架构)下的模拟结果相比,所有配置下预测共享L2缓存缺失率的平均绝对误差小于2%。在与包含相干缺失的上游模型集成后,由于误差累积,4种硬件配置的总体平均绝对误差降至4.82%。作为集成模型的应用案例,我们还评估了57种不同的多核和多级缓存配置的缺失率。
To mitigate the ever worsening “Power wall” and “Memory wall” problems, multi-core architectures with multi-level cache hierarchies have been widely accepted in modern processors. However, the complexity of the architectures makes modeling of shared caches extremely complex. In this article, we propose a data-sharing aware analytical model for estimating the miss rates of the downstream shared cache under multi-core scenarios. To avoid time-consuming full simulations of the cache architecture required by conventional approaches, the proposed model can also be integrated with our refined upstream cache analytical model, which also evaluates coherence misses with similar accuracies of state-of-the-art approach with only one tenth time overhead. We validate our analytical model against gem5 simulation results under 13 applications from PARSEC 2.1 benchmark suites. Compared to the results from gem5 simulations under 8 hardware configurations including dual-core and quad-core architectures, the average absolute error of the predicted shared L2 cache miss rates is less than 2% for all configurations. After integrated with the refined upstream model with coherence misses, the overall average absolute error in 4 hardware configurations is degraded to 4.82% due to the error accumulations. As an application case of the integrated model, we also evaluate the miss rates of 57 different multi-core and multi-level cache configurations.
德克萨斯大学奥斯汀分校/德克萨斯大学奥斯汀分校(美国)
DOI: --
发表时间: --
期刊:
影响因子: --
作者:
通讯作者: --