Isoefficiency in Practice: Configuring and Understanding the Performance of Task-based Applications

Isoefficiency in Practice: Configuring and Understanding the Performance of Task-based Applications
复制标题

实践中的等效率:配置和了解基于任务的应用程序的性能

DOI:
10.1145/3018743.3018770
复制
发表时间:
2017
期刊:
Proceedings of the 22nd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming
影响因子:
--
通讯作者:
Torsten
Torsten
中科院分区:
--
文献类型:
--
作者:
Shudler;Sergei;Calotoiu;Alexandru;Hoefler;Torsten

文献摘要

参考文献

被引文献

相似文献

基于任务的编程提供了一种很好的方式来表达计算单元及其之间的依赖关系,从而更容易在多个内核之间平均分配计算负载。然而,这种问题分解和并行性的分离需要足够大的输入问题,才能在给定数量的核上获得令人满意的效率。不幸的是,在输入大小和内核计数之间找到良好的匹配通常需要大量的实验,这是昂贵的,有时甚至是不切实际的。在本文中,我们提出了一种自动经验方法,用于在一个解析表达式中求出基于任务的程序的等效率函数、绑定效率、核数和输入大小。这使得后两者可以根据给定的(现实的)效率目标进行调整。此外,我们不仅找到(I)实际的等效率函数,而且(Ii)在程序执行没有资源争用的情况下所产生的函数,以及(Iii)只有当程序在整个执行过程中能够保持其平均并行度时才能达到的上限。这三者之间的差异有助于解释低效率,尤其是有助于区分资源争用和与任务依赖或调度相关的结构性冲突。所获得的见解可用于共同设计程序和共享系统资源。
Task-based programming offers an elegant way to express units of computation and the dependencies among them, making it easier to distribute the computational load evenly across multiple cores. However, this separation of problem decomposition and parallelism requires a sufficiently large input problem to achieve satisfactory efficiency on a given number of cores. Unfortunately, finding a good match between input size and core count usually requires significant experimentation, which is expensive and sometimes even impractical. In this paper, we propose an automated empirical method for finding the isoefficiency function of a task-based program, binding efficiency, core count, and the input size in one analytical expression. This allows the latter two to be adjusted according to given (realistic) efficiency objectives. Moreover, we not only find (i) the actual isoefficiency function but also (ii) the function one would yield if the program execution was free of resource contention and (iii) an upper bound that could only be reached if the program was able to maintain its average parallelism throughout its execution. The difference between the three helps to explain low efficiency, and in particular, it helps to differentiate between resource contention and structural conflicts related to task dependencies or scheduling. The insights gained can be used to co-design programs and shared system resources.
使用自动化性能建模来查找复杂代码中的可扩展性错误
DOI: 10.1145/2503210.2503277
发表时间: 2013
期刊: 2013 SC - International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子: --
作者:
Calotoiu;Hoefler
通讯作者: Hoefler
描述和减轻任务并行程序中的工作时间膨胀
DOI: --
发表时间: 2012
期刊: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
Stephen L. Olivier;B. Supinski;M. Schulz;J. Prins
通讯作者: J. Prins
DOI: 10.1145/1413370.1413407
发表时间: 2008-11
期刊: 2008 SC - International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
A. Duran;J. Corbalán;E. Ayguadé
通讯作者: A. Duran;J. Corbalán;E. Ayguadé
用于理解应用程序扩展问题的 MPI 实现的性能模型
DOI: 10.1007/978-3-642-15646-5_3
发表时间: 2010
影响因子: 7.2
作者:
T. Hoefler;W. Gropp;R. Thakur;J. Träff
通讯作者: J. Träff
DOI: 10.1145/2751205.2751216
发表时间: 2015-06
期刊: Proceedings of the 29th ACM on International Conference on Supercomputing
影响因子: --
作者:
Sergei Shudler;A. Calotoiu;T. Hoefler;A. Strube;F. Wolf
通讯作者: Sergei Shudler;A. Calotoiu;T. Hoefler;A. Strube;F. Wolf