Nexus: A New Approach to Replication in Distributed Shared Caches

Nexus: A New Approach to Replication in Distributed Shared Caches
复制标题

Nexus:分布式共享缓存中复制的新方法

DOI:
10.1109/pact.2017.42
复制
发表时间:
2017
期刊:
2017 26th International Conference on Parallel Architectures and Compilation Techniques (PACT)
影响因子:
--
通讯作者:
Daniel Sánchez
Daniel Sánchez
中科院分区:
--
文献类型:
--
作者:
Po;Nathan Beckmann;Daniel Sánchez

文献摘要

被引文献

相似文献

最后一层的缓存越来越分布,由许多小型银行组成。为了表现良好,大多数访问必须由靠近核心的银行提供服务。一种有吸引力的方法是复制仅阅读数据,以便附近有副本。但是,复制在容量和延迟之间引入了微妙的权衡:复制力太少,无法进入遥远的银行,而过多的复制浪费了缓存空间,并导致过度的芯片片段失误。工作负载在所需的复制量方面差异很大,要求采用适应性方法。事先自适应复制技术仅在每个图块的本地银行中复制数据,因此他们专注于选择要复制的数据。不幸的是,未经复制的数据仍然会引起完整的网络遍历,限制了这些技术的性能。我们认为,更好的策略是让核心共享副本,并且自适应方案应专注于选择要复制多少(即多少个)复制品可以穿过芯片)。这个想法充分利用了潜伏能力的权衡,比以前的自适应复制技术在质量上取得更高的性能。它可以应用于许多先前的缓存组织,我们在两个上进行了证明:Nexus-R扩展了R-NUCA,Nexus-J扩展了拼图。我们评估了在HPC上的Nexus和在144核芯片上运行的服务器工作负载,在该芯片上,它的表现优于先前的自适应复制方案,并在所有对复制方面敏感的工作负载中的绩效平均提高了90%,平均提高了23%。
Last-level caches are increasingly distributed, consisting of many small banks. To perform well, most accesses must be served by banks near requesting cores. An attractive approach is to replicate read-only data so that a copy is available nearby. But replication introduces a delicate tradeoff between capacity and latency: too little replication forces cores to access faraway banks, while too much replication wastes cache space and causes excessive off-chip misses. Workloads vary widely in their desired amount of replication, demanding an adaptive approach. Prior adaptive replication techniques only replicate data in each tile's local bank, so they focus on selecting which data to replicate. Unfortunately, data that is not replicated still incurs a full network traversal, limiting the performance of these techniques.We argue that a better strategy is to let cores share replicas and that adaptive schemes should focus on selecting how much to replicate (i.e., how many replicas to have across the chip). This idea fully exploits the latency-capacity tradeoff, achieving qualitatively higher performance than prior adaptive replication techniques. It can be applied to many prior cache organizations, and we demonstrate it on two: Nexus-R extends R-NUCA, and Nexus-J extends Jigsaw. We evaluate Nexus on HPC and server workloads running on a 144-core chip, where it outperforms prior adaptive replication schemes and improves performance by up to 90% and by 23% on average across all workloads sensitive to replication.