Coherence domain restriction on large scale systems

Coherence domain restriction on large scale systems
复制标题

大规模系统的相干域限制

DOI:
--
复制
发表时间:
2015
期刊:
Micro
影响因子:
--
通讯作者:
D. Wentzlaff
D. Wentzlaff
中科院分区:
--
文献类型:
--
作者:
Yaosheng Fu;Tri M. Nguyen;D. Wentzlaff

文献摘要

被引文献

相似文献

设计大规模高速缓存一致性系统一直是一个难以实现的目标。无论是在大型GPU上,还是在未来的千核芯片上,还是在百万核仓库规模的计算机上,共享内存,即使在有限的程度上,都可以提高可编程性。这项工作通过将一致性限制到系统总核心和主节点的灵活子集(域)来避免创建大规模可扩展缓存一致性的传统挑战。提出了一种新的一致性框架--一致性域限制(CDR),该框架能够在保持较低的存储和能量开销的情况下创建使用共享内存的数千到数百万个核心系统。由于有限的应用程序并行性或有限的页面共享,CDR观察到大多数高速缓存线仅由核心的子集共享,CDR将一致性域从全局高速缓存一致性限制到虚拟机级、应用级或页面级。我们探讨了两种类型的限制,一种限制可以访问一致性域的共享者的总数,另一种限制参与一致性域的归属节点的数量和位置。每个独立的相干域只跟踪其域中的核心,而不是整个系统,因此不需要建立在CDR之上的一致性方案来进行扩展。随着核心计数的增加,共享器限制实现了持续的存储开销,而Home限制提供了本地化通信,从而实现了更高的性能。与以前的系统不同,CDR是灵活的,并且不限制归属节点或共享者在域中的位置。我们在1024核芯片的背景下以及共享内存在1,000,000核仓库规模的计算机上的新应用中对CDR进行了评估。共享者限制显著节省了面积,而1024核芯片和1,000,000核系统中的Home限制与全球Home Place相比,性能分别提高了29%和23.04倍。我们在采用IBM的32 nm SOI工艺的25核处理器中实现了整个CDR框架,并给出了详细的区域表征。
Designing massive scale cache coherence systems has been an elusive goal. Whether it be on large-scale GPUs, future thousand-core chips, or across million-core warehouse scale computers, having shared memory, even to a limited extent, improves programmability. This work sidesteps the traditional challenges of creating massively scalable cache coherence by restricting coherence to flexible subsets (domains) of a system's total cores and home nodes. This paper proposes Coherence Domain Restriction (CDR), a novel coherence framework that enables the creation of thousand to million core systems that use shared memory while maintaining low storage and energy overhead. Inspired by the observation that the majority of cache lines are only shared by a subset of cores either due to limited application parallelism or limited page sharing, CDR restricts the coherence domain from global cache coherence to VM-level, application-level, or page-level. We explore two types of restriction, one which limits the total number of sharers that can access a coherence domain and one which limits the number and location of home nodes that partake in a coherence domain. Each independent coherence domain only tracks the cores in its domain instead of the whole system, thereby removing the need for a coherence scheme built on top of CDR to scale. Sharer Restriction achieves constant storage overhead as core count increases while Home Restriction provides localized communication enabling higher performance. Unlike previous systems, CDR is flexible and does not restrict the location of the home nodes or sharers within a domain. We evaluate CDR in the context of a 1024-core chip and in the novel application of shared memory to a 1,000,000-core warehouse scale computer. Sharer Restriction results in significant area savings, while Home Restriction in the 1024-core chip and 1,000,000-core system increases performance by 29% and 23.04× respectively when comparing with global home placement. We implemented the entire CDR framework in a 25-core processor taped out in IBM's 32nm SOI process and present a detailed area characterization.