A shared cache for a chip multi vector processor

A shared cache for a chip multi vector processor
复制标题

一种芯片多向量处理器的共享缓存

DOI:
10.1145/1509084.1509088
复制
发表时间:
2008
期刊:
33rd International Symposium on Computer Architecture (ISCA'06)
影响因子:
--
通讯作者:
Hiroaki Kobayashi
Hiroaki Kobayashi
中科院分区:
--
文献类型:
--
作者:
A. Musa;Y. Sato;Takashi Soga;Koki Okabe;Ryusuke Egawa;Hiroyuki Takizawa;Hiroaki Kobayashi

文献摘要

参考文献

被引文献

相似文献

本文讨论了芯片多矢量处理器(CMVP)的设计,尤其是在芯片内存储器带宽有限时检查芯片缓存的效果。由于芯片多处理器(CMP)已成为商品标量处理器的主流,因此CMP体系结构将在不久的将来设计用于矢量处理器的设计,以利用芯片上的大量晶体管。为了保持更高的科学和工程应用执行持续性能,矢量处理器(CORE)通常需要记忆带宽与至少4个字节/flop(B/Flop)的算术性能之比。但是,由于PIN带宽有限,向量超级计算机已经遇到了存储墙问题。因此,我们提出了一个片上共享的缓存,以维护CMVP的有效内存带宽。我们使用实际的科学应用程序根据NEC SX矢量体系结构评估CMVP的性能。特别是,当B/Flop率降低时,我们检查了对持续性能的缓存效应。实验结果表明,与没有缓存相比,8 MB片共享的缓存可以将四核CMVP的性能提高15%至40%。这是因为共享缓存可以提高多线程的高速缓存率。在这里,共享的缓存采用了MISS状态处理寄存器,该寄存器有可能加速科学和工程应用中的差异方案。此外,我们表明,当使用片上缓存时,2 b/flop足以使CMVP达到高可扩展性。
This paper discusses the design of a chip multi vector processor (CMVP), especially examining the effects of an on-chip cache when the off-chip memory bandwidth is limited. As chip multiprocessors (CMPs) have become the mainstream in commodity scalar processors, the CMP architecture will be adopted to design of vector processors in the near future for harnessing a large number of transistors on a chip. To keep a higher sustained performance in execution of scientific and engineering applications, a vector processor (core) generally requires the ratio of the memory bandwidth to the arithmetic performance of at least 4 bytes/flop (B/FLOP). However, vector supercomputers have been encountering the memory wall problem due to the limited pin bandwidth. Therefore, we propose an on-chip shared cache to maintain the effective memory bandwidth for a CMVP. We evaluate the performance of the CMVP based on the NEC SX vector architecture using real scientific applications. Especially, we examine the caching effect on the sustained performance when the B/FLOP rate is decreased. The experimental results indicate that an 8 MB on-chip shared cache can improve the performance of a four-core CMVP by 15% to 40%, compared with that without the cache. This is because the shared cache can increase cache hit rates of multi-threads. Here, the shared cache employs a miss status handling registers, which has the potential for accelerating difference schemes in scientific and engineering applications. Moreover, we show that the 2 B/FLOP is enough for the CMVP to achieve a high scalability when the on-chip cache is employed.
矢量处理器的片上高速缓存设计
DOI: --
发表时间: 2007
期刊: Proceedings of the 8th MEDEA workshop
影响因子: --
作者:
Akihiro Musa;Yoshiei Sato;Ryusuke Egawa;Hiroyuki Takizawa;Koki Okabe and Hiroaki Kobayashi
通讯作者: Koki Okabe and Hiroaki Kobayashi
DOI: --
发表时间: 2006
期刊: Proceedings of International Symposium on Parallel and Distributed Processing and Application (ISPA06)
影响因子: --
作者:
Akihiko Musa;Hiroyuki Takizawa;Koki Okabe;Takashi Soga and Hiroaki Kobayashi
通讯作者: Takashi Soga and Hiroaki Kobayashi
HEC 系统中内存性能的影响
DOI: --
发表时间: 2006
期刊: High Performance Computing on Vector Systems(Springer-Verlag)
影响因子: --
作者:
Yukinori Sato;Ken-ichi Suzuki and Tadao Nakamura;Hiroaki Kobayashi
通讯作者: Hiroaki Kobayashi