Multigrain parallel processing on compiler cooperative chip multiprocessor

Multigrain parallel processing on compiler cooperative chip multiprocessor
复制标题

编译器协同芯片多处理器的多粒度并行处理

DOI:
10.1109/interact.2005.9
复制
发表时间:
2005
期刊:
9th Annual Workshop on Interaction between Compilers and Computer Architectures (INTERACT'05)
影响因子:
--
通讯作者:
H. Kasahara
H. Kasahara
中科院分区:
--
文献类型:
--
作者:
K. Kimura;Y. Wada;H. Nakano;T. Kodaka;J. Shirako;K. Ishizaka;H. Kasahara

文献摘要

参考文献

被引文献

相似文献

本文描述了一种编译器协同片上多处理机的多粒度并行处理。多粒度并行处理分层地利用多粒度并行,例如粗粒度任务并行、循环迭代级并行和语句级近细粒度并行。片上多处理器是通过支持日本千年计划IT21“Advance并行编译器”的多粒度并行编译器的优化来实现高效性能、成本效益和高软件生产率的。为了实现多粒度并行处理的全部潜力,芯片多处理器集成了简单的单问题处理器,其具有用于数据局部性和标量数据传输的最佳使用的分布式共享数据存储器,用于处理器私有数据的本地数据存储器,以及用于处理器之间共享数据的集中式共享存储器。本文重点研究了利用SPECfp95程序的多粒度并行性,在一个片上具有多达8个处理器的片上多处理器的可扩展性。在90 nm工艺和2.8 GHz的假设下使用简单处理器内核等微处理器时,评估结果显示,8个处理器和4个处理器的加速比分别达到7.1和3.9。类似地,当400 MHz被假定用于嵌入式使用时,加速比分别达到7.8和4.0。
This paper describes multigrain parallel processing on a compiler cooperative chip multiprocessor. The multigrain parallel processing hierarchically exploits multiple grains of parallelism such as coarse grain task parallelism, loop iteration level parallelism and statement level near-fine grain parallelism. The chip multiprocessor has been designed to attain high effective performance, cost effectiveness and high software productivity by supporting the optimizations of the multigrain parallelizing compiler, which is developed by Japanese Millennium Project IT21 "Advance Parallelizing Compiler". To achieve full potential of multigrain parallel processing, the chip multiprocessor integrates simple single-issue processors having distributed shared data memory for both optimal use of data locality and scalar data transfer, local data memory for processor private data, in addition to centralized shared memory for shared data among processors. This paper focuses on the scalability of the chip multiprocessor having up to eight processors on a chip by exploiting of the multigrain parallelism from SPECfp95 programs. When microSPARC like the simple processor core is used under assumption of 90 nm technology and 2.8 GHz, the evaluation results show the speedups for eight processors and four processors reach 7.1 and 3.9, respectively. Similarly, when 400 MHz is assumed for embedded usage, the speedups reach 7.8 and 4.0, respectively.
DOI: 10.11203/jar.19.177
发表时间: 2004-09
期刊: --
影响因子: --
作者:
裕幸 飯田;菊男 竹田;武利 藤本
通讯作者: 裕幸 飯田;菊男 竹田;武利 藤本