Analysis of cache invalidation patterns in multiprocessors

Analysis of cache invalidation patterns in multiprocessors
复制标题

多处理器中缓存失效模式分析

DOI:
--
复制
发表时间:
1989
期刊:
ASPLOS III
影响因子:
--
通讯作者:
Anoop Gupta
Anoop Gupta
中科院分区:
--
文献类型:
--
作者:
W. Weber;Anoop Gupta

文献摘要

被引文献

相似文献

为了使共享内存多处理器可扩展,研究人员现在正在探索不依赖于广播的缓存一致性协议,而是向包含陈旧数据的各个缓存发送无效消息。这种基于目录的协议的可行性对并行程序所表现出的该高速缓存失效模式高度敏感。在本文中,我们分析了该高速缓存失效模式所造成的几个并行应用程序,并调查这些模式的影响,基于目录的协议。我们的结果是基于多处理器的痕迹与4,8和16个处理器。为了深入了解失效模式会看起来像超过16个处理器,我们提出了一个分类方案,在并行应用程序中发现的数据对象,并链接到这些高级别对象的痕迹中观察到的失效流量模式。我们的研究结果表明,同步对象有非常不同的失效模式从其他数据对象。对同步对象的写引用通常会导致更多缓存中的无效。我们指出的情况下,重组的应用程序似乎是适当的,以减少无效流量,和其他的硬件支持是更合适的。我们的研究结果还表明,它应该是可能的规模“写得很好”的并行程序,以大量的处理器,而不会在无效流量爆炸。
To make shared-memory multiprocessors scalable, researchers are now exploring cache coherence protocols that do not rely on broadcast, but instead send invalidation messages to individual caches that contain stale data. The feasibility of such directory-based protocols is highly sensitive to the cache invalidation patterns that parallel programs exhibit. In this paper, we analyze the cache invalidation patterns caused by several parallel applications and investigate the effect of these patterns on a directory-based protocol. Our results are based on multiprocessor traces with 4, 8 and 16 processors. To gain insight into what the invalidation patterns would look like beyond 16 processors, we propose a classification scheme for data objects found in parallel applications and link the invalidation traffic patterns observed in the traces back to these high-level objects. Our results show that synchronization objects have very different invalidation patterns from those of other data objects. A write reference to a synchronization object usually causes invalidations in many more caches. We point out situations where restructuring the application seems appropriate to reduce the invalidation traffic, and others where hardware support is more appropriate. Our results also show that it should be possible to scale “well-written” parallel programs to a large number of processors without an explosion in invalidation traffic.