Hierarchical Cache Directory for CMP

Hierarchical Cache Directory for CMP
复制标题

DOI:
10.1007/s11390-010-9321-5
复制
发表时间:
2010-03
影响因子:
0.7
通讯作者:
Songliu Guo;Haixia Wang;Y. Xue;Chongmin Li;Dong-Sheng Wang
Songliu Guo;Haixia Wang;Y. Xue;Chongmin Li;Dong-Sheng Wang
中科院分区:
--
文献类型:
--
作者:
Songliu Guo;Haixia Wang;Y. Xue;Chongmin Li;Dong-Sheng Wang

文献摘要

被引文献

相似文献

随着更多的处理核心集成到一个芯片中,并且特征尺寸不断缩小,使用基于目录的一致性协议的远程节点的平均访问延迟变得更高,这极大地影响了系统性能。以前的技术,如数据复制和数据迁移,优化性能的请求核心,但提供的邻居节点的改进很少。其他技术(如传输中优化)试图以增加存储为代价来减少延迟。本文将层次化的Cache目录引入片上多处理器中,将片上多处理器分片分层,并将其与数据复制相结合。提出了一种新的目录组织结构,用于记录区域内的共享状态,辅助区域归属地高效地完成操作。仿真结果表明,对于16核CMP,与传统目录相比,分层缓存目录在存储空间较少的情况下,平均访问延迟降低了9%,片上网络流量平均降低了34%.理论分析表明,对于一个2n× 2n分片的CMP,层次缓存目录的平均访问延迟渐近地接近于一个与n无关的函数,因此该体系结构具有很强的可扩展性.
As more processing cores are integrated into one chip and feature size continues to shrink, the average access latency for remote nodes using directory-based coherence protocol becomes higher, which greatly impacts system performance. Previous techniques such as data replication and data migration optimize the performance of the requesting core, but offer little improvement for neighbor nodes. Other techniques such as in-transit optimization try to reduce latency at the cost of increased storage. This paper introduces hierarchical cache directory into CMP (chip multiprocessor), which divides CMP tiles into multiple regions hierarchically, and combines it with data replication. A new directory organization is proposed to record the share status within a region and assist the regional home to complete operation efficiently. Simulation results show that for a 16-core CMP, compared to traditional directory, hierarchical cache directory reduces average access latency by 9% and on-chip network traffic by 34% on average with less storage. Theoretical analyses show that for a 2n× 2ntiled CMP, the average access latency in hierarchical cache directory asymptotically approaches a function that is independent ofn, hence the architecture is highly scalable.