On the value locality of store instructions

On the value locality of store instructions
复制标题

DOI:
10.1145/339647.339678
复制
发表时间:
2000-05
期刊:
Proceedings of 27th International Symposium on Computer Architecture (IEEE Cat. No.RS00201)
影响因子:
--
通讯作者:
Kevin M. Lepak;Mikko H. Lipasti
Kevin M. Lepak;Mikko H. Lipasti
中科院分区:
其他
文献类型:
--
作者:
Kevin M. Lepak;Mikko H. Lipasti

文献摘要

被引文献

相似文献

价值局部性是最近发现的一种程序属性,它描述了以前看到的程序值重复出现的可能性,在最近发表的文献中得到了热烈的研究。大部分精力都集中在改进预测负载指令结果的初始工作上,同时检查所有寄存器写入指令或它们的一个集中子集的值位置。令人惊讶的是,很少有关于计算机程序存储在存储器中的数据字的值局部性的描述或研究。本文提出了这样一个特征,提出了存储数据值的以内存为中心(基于消息传递)和以生产者为中心(基于程序结构)的预测机制,介绍了静默存储的概念和基于这些观察的多处理器错误共享的新定义,并提出了对齐缓存一致性协议和微架构存储处理技术的新技术,以利用存储的值局域性。我们发现这些技术的实际实现可以显着减少多处理器数据总线流量,并且在减少地址总线流量方面比在MSI一致性协议中添加独占状态更有效。我们还表明,压缩静默存储可以提供比添加存储到负载转发更大的单处理器速度。
Value locality, a recently discovered program attribute that describes the likelihood of the recurrence of previously-seen program values, has been studied enthusiastically in the recent published literature. Much of the energy has focused on refining the initial efforts at predicting load instruction outcomes, with the balance of the effort examining the value locality of either all register-writing instructions, or a focused subset of them. Surprisingly, there has been very little published characterization of or effort to exploit the value locality of data words stored to memory by computer programs. This paper presents such a characterization, proposes both memory-centric (based on message passing) and producer-centric (based on program structure) prediction mechanisms for stored data values, introduces the concept of silent stores and new definitions of multiprocessor false sharing based on these observations, and suggests new techniques for aligning cache coherence protocols and microarchitectural store handling techniques to exploit the value locality of stores. We find that realistic implementations of these techniques can significantly reduce multiprocessor data bus traffic and are more effective at reducing address bus traffic than the addition of Exclusive state to a MSI coherence protocol. We also show that squashing of silent stores can provide uniprocessor speedups greater than the addition of store-to-load forwarding.