Cache performance in vector supercomputers

Cache performance in vector supercomputers
复制标题

矢量超级计算机中的缓存性能

DOI:
10.1109/superc.1994.344285
复制
发表时间:
1994
期刊:
Proceedings of Supercomputing '94
影响因子:
--
通讯作者:
M. Scott
M. Scott
中科院分区:
--
文献类型:
--
作者:
L. Kontothanassis;R. Sugumar;Greg Faanes;James E. Smith;M. Scott

文献摘要

被引文献

相似文献

传统的超级计算机使用平坦的多银行SRAM内存组织在低潜伏期下提供高带宽。大多数其他计算机都使用带有小型SRAM缓存的分层组织,用于主内存的较小,更便宜的DRAM。这样的系统在很大程度上依赖数据局部性来实现最佳性能。本文评估了用于向量超级计算机的基于缓存的内存系统。我们为基于缓存的Cray Research C90的基于缓存的版本开发了模拟模型,并使用NAS并行基准测试提供了大规模的工作量。我们表明,虽然缓存减少记忆流量并改善普通DRAM内存的性能,但它们仍然落后于Cacheless SRAM。我们确定基于DRAM的内存系统中的性能瓶颈,并量化其对程序性能降解的贡献。我们发现数据获取策略是影响性能的重要参数,我们评估了几种提取策略的性能,并且我们表明,小提取尺寸通过最大化可用内存带宽的使用来改善性能。<< etx >>
Traditional supercomputers use a flat multi-bank SRAM memory organization to supply high bandwidth at low latency. Most other computers use a hierarchical organization with a small SRAM cache and a slower, cheaper DRAM for the main memory. Such systems rely heavily on data locality for achieving optimum performance. This paper evaluates cache-based memory systems for vector supercomputers. We develop a simulation model for a cache-based version of the Cray Research C90 and use the NAS parallel benchmarks to provide a large-scale workload. We show that while caches reduce memory traffic and improve the performance of plain DRAM memory, they still lag behind cacheless SRAM. We identify the performance bottlenecks in DRAM-based memory systems and quantify their contribution to program performance degradation. We find the data fetch strategy to be a significant parameter affecting performance, we evaluate the performance of several fetch policies, and we show that small fetch sizes improve performance by maximizing the use of available memory bandwidth.<<ETX>>