Performance and Scalability Analysis of a Chip Multi Vector Processor

Performance and Scalability Analysis of a Chip Multi Vector Processor
复制标题

片上多向量处理器的性能和可扩展性分析

DOI:
10.1007/978-3-642-22244-3_1
复制
发表时间:
2012
期刊:
High Performance Computing on Vector Systems 2011
影响因子:
--
通讯作者:
Hiroaki Kobayashi
Hiroaki Kobayashi
中科院分区:
--
文献类型:
--
作者:
Yoshiei Sato;Akihiro Musa;Ryusuke Egawa;Hiroyuki Takizawa;Koki Okabe;Hiroaki Kobayashi

文献摘要

参考文献

相似文献

为了在矢量处理器上实现更高效、更强大的计算能力,提出了一种芯片多矢量处理器(CMVP)作为下一代矢量处理器。然而,CMVP在科学应用中的有用性尚不清楚。本文的目的是阐明CMVP的潜力。虽然CMVP的计算性能随着内核数量的增加而增加,但内存带宽与计算性能的比值(B/F)会下降。为了弥补不足的B/F, CMVP有一个共享的矢量缓存。因此,为了充分发挥CMVP的潜力,不仅需要利用传统的调优技术来提高矢量运算的效率,还需要利用新技术来有效地利用矢量缓存。在这种情况下,本文提出了一种CMVP的性能调优策略。该策略分析应用程序的性能瓶颈,以找到调优技术的最佳组合。通过使用实际应用程序评估调优策略带来的性能和可伸缩性改进。评估结果表明,随着内核数量的增加,性能调优变得更加重要。
To realize more efficient and powerful computations on a vector processor, a chip multi vector processor (CMVP) has been proposed as a next generation vector processor. However, the usefulness of CMVP for scientific applications has been unclear. The objective of this paper is to clarify the potential of CMVP. Although the computational performance of CMVP increases with the number of cores, the ratio of memory bandwidth to computational performance (B/F) will decrease. To cover the insufficient B/F, CMVP has a shared vector cache. Therefore, to exploit the potential of CMVP, applications for CMVP should be optimized not only with conventional tuning techniques to improve the efficiency of vector operations, but also with new techniques to effectively use the vector cache. Under this situation, this paper presents a performance tuning strategy for CMVP. The strategy analyzes the performance bottleneck of an application to find the best combination of tuning techniques. The performance and scalability improvements due to the tuning strategy are evaluated using real applications. The evaluation results clarify that performance tuning becomes more important as the number of cores increases.
基于Roofline模型的未来矢量处理器性能调优与分析
DOI: --
发表时间: 2009
期刊:
影响因子: --
作者:
Yoshiei Sato;Ryuichi Nagaoka;Akihiro Musa;Ryusuke Egawa;Hiroyuki Takizawa;Koki Okabe;Hiroaki Kobayashi
通讯作者: Hiroaki Kobayashi
一种芯片多向量处理器的共享缓存
DOI: 10.1145/1509084.1509088
发表时间: 2008
期刊: 33rd International Symposium on Computer Architecture (ISCA'06)
影响因子: --
作者:
A. Musa;Y. Sato;Takashi Soga;Koki Okabe;Ryusuke Egawa;Hiroyuki Takizawa;Hiroaki Kobayashi
通讯作者: Hiroaki Kobayashi
标量和矢量 HEC 平台上的领先计算方法
DOI: 10.1109/sc.2005.41
发表时间: 2005
期刊: ACM/IEEE SC 2005 Conference (SC'05)
影响因子: --
作者:
L. Oliker;J. Carter;M. Wehner;A. Canning;S. Ethier;A. Mirin;David Parks;P. Worley;S. Kitawaki;Yoshinori Tsuda
通讯作者: Yoshinori Tsuda
片上存储器系统在未来矢量架构中的潜力
DOI: 10.1007/978-3-540-74384-2_18
发表时间: 2008
期刊: 33rd International Symposium on Computer Architecture (ISCA'06)
影响因子: --
作者:
Hiroaki Kobayashi;A. Musa;Y. Sato;Hiroyuki Takizawa;Koki Okabe
通讯作者: Koki Okabe
湍流通道流中物质线的演变
DOI: --
发表时间: 2007
期刊:
影响因子: --
作者:
Tsukahara;T.;Iwamoto;K. and Kawamura;H.
通讯作者: H.