Uncorq: Unconstrained Snoop Request Delivery in Embedded-Ring Multiprocessors

Uncorq: Unconstrained Snoop Request Delivery in Embedded-Ring Multiprocessors
复制标题

Uncorq:嵌入式环多处理器中无约束的侦听请求传送

DOI:
10.1109/micro.2007.43
复制
发表时间:
2007
期刊:
40th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO 2007)
影响因子:
--
通讯作者:
Josep Torrellas
Josep Torrellas
中科院分区:
--
文献类型:
--
作者:
Karin Strauss;Xiaowei Shen;Josep Torrellas

文献摘要

被引文献

相似文献

可以通过在网络中嵌入逻辑单向环来在任何物理网络拓扑中实现Snoopy Cache连贯性。控制消息是使用戒指转发的,而其他消息可以使用任何路径。尽管所得的连贯协议实施便宜,但它们启用了多种方式,将多种交易重叠,这些交易访问相同的线路制造,难以推断正确性。此外,需要侦听请求才能穿越环,从而延长连贯交易潜伏期。在本文中,我们解决了这些问题,并做出了两个主要贡献。首先,我们介绍了排序不变式,该订单确保了嵌入式环协议中碰撞交易的正确序列化。其次,基于此不变性,我们删除了Snoop请求穿越环的要求。取而代之的是,只要使用任何网络路径(通常是关键路径)使用逻辑环,它们就会使用任何网络路径传递。这种方法大大减少了连贯交易延迟。我们称结果协议UNCORQ。我们表明,在64个节点芯片多处理器(CMP)上,UNCORQ平均可以提高splash-2应用程序的性能23%,而商业应用则提高了10%。借助其他简单的预取优化,Splash-2应用程序的性能提高平均为26%,商业应用程序的绩效提高为18%。
Snoopy cache coherence can be implemented in any physical network topology by embedding a logical unidirectional ring in the network. Control messages are forwarded using the ring, while other messages can use any path. While the resulting coherence protocols are inexpensive to implement, they enable many ways of overlapping multiple transactions that access the same line-making it hard to reason about correctness. Moreover, snoop requests are required to traverse the ring, therefore lengthening coherence transaction latencies. In this paper, we address these problems and make two main contributions. First, we introduce the ordering invariant, which ensures the correct serialization of colliding transactions in embedded-ring protocols. Second, based on this invariant, we remove the requirement that snoop requests traverse the ring. Instead, they are delivered using any network path, as long as snoop responses - which are typically off the critical path - use the logical ring. This approach substantially reduces coherence transaction latency. We call the resulting protocol Uncorq. We show that, on a 64-node chip multiprocessor (CMP), Uncorq improves the performance, on average, by 23% for SPLASH-2 applications and by 10% for commercial applications. With an additional simple prefetching optimization, the performance improvement is, on average, 26% for SPLASH-2 applications and 18% for commercial applications.