Callback: Efficient synchronization without invalidation with a directory just for spin-waiting

Callback: Efficient synchronization without invalidation with a directory just for spin-waiting
复制标题

回调:高效同步,无需目录失效,仅用于旋转等待

DOI:
10.1145/2749469.2750405
复制
发表时间:
2015
期刊:
2015 ACM/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
S. Kaxiras
S. Kaxiras
中科院分区:
--
文献类型:
--
作者:
Alberto Ros;S. Kaxiras

文献摘要

被引文献

相似文献

与传统的基于失效的协议相比,基于自失效的缓存一致性协议通过依赖数据无争用 (DRF) 语义并对暴露给硬件的活跃同步点应用自失效,可以实现更简单的设计。它们的简单性在于不存在失效流量,从而无需跟踪目录中的读者,并减少了瞬态协议状态的数量。通过添加自我降级功能,这些协议可以有效地变得无目录。虽然这对于无竞争数据很有效,但不幸的是,缺乏显式失效会损害任何依赖竞争的同步的有效性。这包括用于信令、锁定和屏障原语的任何形式的自旋等待。在这项工作中,我们提出了一种新的协议中的自旋等待解决方案,即回调机制,它比显式失效更简单、更高效。回调由旋转等待中涉及的读取设置,并由写入(甚至可以在这些读取之前)满足。为了实现回调,我们使用一个小型(只有几个条目)目录缓存结构,该结构旨在仅为这些“旋转等待”竞争提供服务。该目录结构是独立的,不以任何方式进行备份。条目是根据需要创建的,并且可以被驱逐而无需保留其信息。我们的评估显示,显式失效和指数退避(自失效协议的最先进机制,以避免共享缓存中的旋转)都有显着改进。
Cache coherence protocols based on self-invalidation allow a simpler design compared to traditional invalidation-based protocols, by relying on data-race-free (DRF) semantics and applying self-invalidation on racy synchronization points exposed to the hardware. Their simplicity lies in the absence of invalidation traffic, which eliminates the need to track readers in a directory, and reduces the number of transient protocol states. With the addition of self-downgrade these protocols can become effectively directory-free. While this works well for race-free data, unfortunately, lack of explicit invalidations compromises the effectiveness of any synchronization that relies on races. This includes any form of spin waiting, which is employed for signaling, locking, and barrier primitives. In this work we propose a new solution for spin-waiting in these protocols, the callback mechanism, that is simpler and more efficient than explicit invalidation. Callbacks are set by reads involved in spin waiting, and are satisfied by writes (that can even precede these reads). To implement callbacks we use a small (just a few entries) directory-cache structure that is intended to service only these “spin-waiting” races. This directory structure is self-contained and is not backed up in any way. Entries are created on demand and can be evicted without the need to preserve their information. Our evaluation shows a significant improvement both over explicit invalidation and over exponential back-off, the state-of-the-art mechanism for self-invalidation protocols to avoid spinning in the shared cache.