Cache Refill/Access Decoupling for Vector Machines

Cache Refill/Access Decoupling for Vector Machines
复制标题

向量机的缓存填充/访问解耦

DOI:
--
复制
发表时间:
2004
期刊:
Micro
影响因子:
--
通讯作者:
K. Asanović
K. Asanović
中科院分区:
--
文献类型:
--
作者:
C. Batten;R. Krashinsky;S. Gerding;K. Asanović

文献摘要

被引文献

相似文献

向量处理器通常使用缓存来利用时间区域并减少内存带宽的需求,但随后需要昂贵的逻辑来跟踪大量出色的高速缓存误差以维持内存的峰值带宽。我们提出了补充/访问解耦合,该解耦合将使用矢量补充单元(VRU)增强矢量处理器,以快速预先执行矢量内存命令,并在常规执行之前发布任何必要的高速缓存线补充。 VRU通过消除传统矢量体系结构中所需的许多出色状态并将缓存本身用作具有成本效益的预购缓冲液,从而降低了成本。我们还介绍了矢量段访问,这是一种有效编码二维访问模式的新类矢量内存指令。段减少地址带宽需求,并通过增加每个向量内存命令中包含的信息来实现更有效的补充/访问解耦。我们的结果表明,与更多传统的解耦方法相比,补充/访问解耦能够通过更少的资源实现更好的性能。即使较小的缓存和记忆潜伏期长达800个周期,补充/访问解耦合也可以维持几千键的机上数据,并且具有最小的访问管理状态,并且不需要昂贵的保留元素缓冲。
Vector processors often use a cache to exploit temporal locality and reduce memory bandwidth demands, but then require expensive logic to track large numbers of outstanding cache misses to sustain peak bandwidth from memory. We present refill/access decoupling, which augments the vector processor with a Vector Refill Unit (VRU) to quickly pre-execute vector memory commands and issue any needed cache line refills ahead of regular execution. The VRU reduces costs by eliminating much of the outstanding miss state required in traditional vector architectures and by using the cache itself as a cost-effective prefetch buffer. We also introduce vector segment accesses, a new class of vector memory instructions that efficiently encode two-dimensional access patterns. Segments reduce address bandwidth demands and enable more efficient refill/access decoupling by increasing the information contained in each vector memory command. Our results show that refill/access decoupling is able to achieve better performance with less resources than more traditional decoupling methods. Even with a small cache and memory latencies as long as 800 cycles, refill/access decoupling can sustain several kilobytes of in-flight data with minimal access management state and no need for expensive reserved element buffering.