Efficiently speeding up sequential computation through the n-way programming model

Efficiently speeding up sequential computation through the n-way programming model
复制标题

通过n路编程模型有效加速顺序计算

DOI:
--
复制
发表时间:
2011
期刊:
Conference on Object-Oriented Programming Systems, Languages, and Applications
影响因子:
--
通讯作者:
S. Pande
S. Pande
中科院分区:
--
文献类型:
--
作者:
R. Cledat;T. Kumar;S. Pande

文献摘要

被引文献

相似文献

随着核心数量的增加,应用程序的顺序组件正成为性能扩展的主要瓶颈,正如Amdahl定律所预测的那样。因此,我们同时面临着占用越来越多的核心和加速连续部分的问题。在这项工作中,我们调和这两个看似不兼容的问题与一种新的编程模型称为N路。N-way背后的核心思想是受益于可用于表达某些关键计算步骤的算法多样性。通过同时并行启动多种方法来解决给定的计算,运行时可以及时选择最佳(例如最快)的方法,从而实现加速。 以前的工作已经证明了这种方法的好处,但没有解决其固有的浪费。在这项工作中,我们专注于提供一个数学上合理的基于学习的统计模型,可以由运行时使用,以确定使用的资源和通过N路可获得的利益之间的最佳平衡。我们进一步描述了一个动态的剔除机制,以进一步减少资源浪费。 我们提出了抽象和运行时支持干净地封装计算选项和监控其进度。我们展示了一个低开销的运行时,实现了显着的加速范围内广泛使用的内核。我们的研究结果表明,在某些情况下,超线性加速。
With core counts on the rise, the sequential components of applications are becoming the major bottleneck in performance scaling as predicted by Amdahl's law. We are therefore faced with the simultaneous problems of occupying an increasing number of cores and speeding up sequential sections. In this work, we reconcile these two seemingly incompatible problems with a novel programming model called N-way. The core idea behind N-way is to benefit from the algorithmic diversity available to express certain key computational steps. By simultaneously launching in parallel multiple ways to solve a given computation, a runtime can just-in-time pick the best (for example the fastest) way and therefore achieve speedup. Previous work has demonstrated the benefits of such an approach but has not addressed its inherent waste. In this work, we focus on providing a mathematically sound learning-based statistical model that can be used by a runtime to determine the optimal balance between resources used and benefits obtainable through N-way. We further describe a dynamic culling mechanism to further reduce resource waste. We present abstractions and a runtime support to cleanly encapsulate the computational-options and monitor their progress. We demonstrate a low-overhead runtime that achieves significant speedup over a range of widely used kernels. Our results demonstrate super-linear speedups in certain cases.