Conflict-free accesses to strided vectors on a banked cache

Conflict-free accesses to strided vectors on a banked cache
复制标题

对存储缓存上的跨步向量进行无冲突访问

DOI:
10.1109/tc.2005.110
复制
发表时间:
2005
影响因子:
3.7
通讯作者:
R. Espasa
R. Espasa
中科院分区:
计算机科学2区
文献类型:
--
作者:
André Seznec;R. Espasa

文献摘要

被引文献

相似文献

随着集成技术的进步,实现微处理器,矢量单元和一个多层银行交换的L2 Cache在单个模具上变得可行。在L2高速缓存上对跨性别向量的并行访问是此类矢量微处理器的主要性能问题。这样的并行访问的主要困难是,人们希望以块大小为基础交织缓存,以便从空间区域中受益并保持低标签量,而障碍的矢量访问自然可以在单词粒度上工作。在本文中,我们解决了这个问题。考虑一个平行矢量单元,具有2/sup n/独立车道,2/sup n/bank交织高速缓存以及2/sup k/单词的高速缓存线大小,我们表明任何2/sup n+k/sup n+k/sup k/sup k/单词在L2缓存中可以访问任何具有r奇数和r/spl les/k的跨度2/sup r/r的跨载体的连续元素,并以2/sup n/sup n/sup -subsines中的2/sup k/subsline返回泳道元素。
With the advance of integration technology, it has become feasible to implement a microprocessor, a vector unit, and a multimegabyte bank-interleaved L2 cache on a single die. Parallel access to strided vectors on the L2 cache is a major performance issue on such vector microprocessors. A major difficulty for such a parallel access is that one would like to interleave the cache on a block size basis in order to benefit from spatial locality and to maintain a low tag volume, while strided vector accesses naturally work on a word granularity. In this paper, we address this issue. Considering a parallel vector unit with 2/sup n/ independent lanes, a 2/sup n/ bank interleaved cache, and a cache line size of 2/sup k/ words, we show that any slice of 2/sup n+k/ consecutive elements of any strided vector with stride 2/sup r/R with R odd and r /spl les/ k can be accessed in the L2 cache and routed back to the lanes in 2/sup k/ subslices of 2/sup n/ elements.