Loop Parallelism: A New Skeleton Perspective on Data Parallel Patterns

Loop Parallelism: A New Skeleton Perspective on Data Parallel Patterns
复制标题

循环并行:数据并行模式的新骨架视角

DOI:
--
复制
发表时间:
2014
期刊:
2014 22nd Euromicro International Conference on Parallel, Distributed, and Network-Based Processing
影响因子:
--
通讯作者:
M. Torquati
M. Torquati
中科院分区:
--
文献类型:
--
作者:
M. Danelutto;M. Torquati

文献摘要

被引文献

相似文献

传统上,基于骨架的并行编程框架通过为程序员提供基于MAP的不同变体和减少模式的全面数据并行骨骼来支持数据并行性。另一方面,更传统的并行编程框架为应用程序程序员提供了可以在循环执行循环中以相对较小的编程工作来引入并行性的。在这项工作中,我们讨论了在快速流框架内提供的“同行”骨架,旨在填补经典数据并行骨架方法与诸如OpenMP和Intel TBB等框架提供的环路并行设施之间的可用性和表达差距。通过利用快速流平行骨架的低运行时间开销以及C ++ 11标准提供的新设施,我们的同行骨骼成功地获得了与Intel Phi多核和Intel上的OpenMP和TBB相比的可比性或更好的性能Nehalem多核用于考虑一组基准测试,但需要相当的编程工作。
Traditionally, skeleton based parallel programming frameworks support data parallelism by providing the programmer with a comprehensive set of data parallel skeletons, based on different variants of map and reduce patterns. On the other side, more conventional parallel programming frameworks provide application programmers with the possibility to introduce parallelism in the execution of loops with a relatively small programming effort. In this work, we discuss a "ParallelFor" skeleton provided within the FastFlow framework and aimed at filling the usability and expressivity gap between the classical data parallel skeleton approach and the loop parallelisation facilities offered by frameworks such as OpenMP and Intel TBB. By exploiting the low run-time overhead of the FastFlow parallel skeletons and the new facilities offered by the C++11 standard, our ParallelFor skeleton succeeds to obtain comparable or better performance than both OpenMP and TBB on the Intel Phi many-core and Intel Nehalem multi-core for a set of benchmarks considered, yet requiring a comparable programming effort.