The effect of sharing on the cache and bus performance of parallel programs

The effect of sharing on the cache and bus performance of parallel programs
复制标题

共享对并行程序的缓存和总线性能的影响

DOI:
--
复制
发表时间:
1989
期刊:
ASPLOS III
影响因子:
--
通讯作者:
R. Katz
R. Katz
中科院分区:
--
文献类型:
--
作者:
S. Eggers;R. Katz

文献摘要

被引文献

相似文献

总线带宽最终限制了基于总线的共享内存多处理器的性能,因此也限制了其规模。以前的研究已经从单处理器的测量和模拟来推断这些机器的性能。在这项研究中,我们使用的并行程序的痕迹来评估该高速缓存和总线性能的共享内存多处理器,其中一致性是由写无效协议保持。特别是,我们分析了共享开销对高速缓存未命中率和总线利用率的影响。 我们的研究表明,并行程序比同类单处理器程序产生更高的未命中率和总线利用率。这些指标的共享部分随着缓存和数据块大小成比例地增加,并且对于某些缓存配置,决定了它们的大小和趋势。开销量取决于对共享数据的内存引用模式。表现出良好的每处理器局部性的程序比细粒度共享的程序性能更好。这表明并行软件编写者和更好的编译器技术可以通过更好的共享数据内存组织来提高程序性能。
Bus bandwidth ultimately limits the performance, and therefore the scale, of bus-based, shared memory multiprocessors. Previous studies have extrapolated from uniprocessor measurements and simulations to estimate the performance of these machines. In this study, we use traces of parallel programs to evaluate the cache and bus performance of shared memory multiprocessors, in which coherency is maintained by a write-invalidate protocol. In particular, we analyze the effect of sharing overhead on cache miss ratio and bus utilization. Our studies show that parallel programs incur substantially higher miss ratios and bus utilization than comparable uniprocessor programs. The sharing component of these metrics proportionally increases with both cache and block size, and for some cache configurations determines both their magnitude and trend. The amount of overhead depends on the memory reference pattern to the shared data. Programs that exhibit good per-processor-locality perform better than those with fine-grain-sharing. This suggests that parallel software writers and better compiler technology can improve program performance through better memory organization of shared data.