Accelerating decoupled look-ahead via weak dependence removal: A metaheuristic approach

Accelerating decoupled look-ahead via weak dependence removal: A metaheuristic approach
复制标题

通过弱依赖消除加速解耦前瞻:一种元启发式方法

DOI:
--
复制
发表时间:
2014
期刊:
International Symposium on High-Performance Computer Architecture
影响因子:
--
通讯作者:
Michael C. Huang
Michael C. Huang
中科院分区:
--
文献类型:
--
作者:
Raj Parihar;Michael C. Huang

文献摘要

被引文献

相似文献

尽管多核和多线程架构的扩散,但对单个语义线程的隐式并行性仍然是实现高性能的关键组成部分。 ,在远距离远距离时,单层的核心核心很快就会成为资源感知。多核环境能够产生大量的性能,但往来的速度限制通常是新的速度限制。 “弱”的依赖性,并非所有依赖性依赖性链中的某些链接。常见的依赖性模式,因为为了生成更好的代码,它们的启发式是一个主要的原因。确定基于遗传算法的框架的机会,可以帮助搜索适用于look-ead线程的框架。新的限制是,这种方法可将整体系统性能提高1.48倍,而几何平均值在基线解耦的位置系统上为1.14倍,同时将能源消耗降低了11%。
Despite the proliferation of multi-core and multi-threaded architectures, exploiting implicit parallelism for a single semantic thread is still a crucial component in achieving high performance. Look-ahead is a tried-and-true strategy in uncovering implicit parallelism, but a conventional, monolithic out-of-order core quickly becomes resource-inefficient when looking beyond a small distance. A more decoupled approach with an independent, dedicated look-ahead thread on a separate thread context can be a more flexible and effective implementation, especially in a multi-core environment. While capable of generating significant performance gains, the look-ahead agent often becomes the new speed limit. Fortunately, the look-ahead thread has no hard correctness constraints and presents new opportunities for optimizations. One such opportunity is to exploit “weak” dependences. Intuitively, not all dependences are equal. Some links in a dependence chain are weak enough that removing them in the look-ahead thread does not materially affect the quality of look-ahead but improves the speed. While there are some common patterns of weak dependences, they can not be generalized as heuristics in generating better code for the look-ahead thread. A primary reason is that removing a false weak dependence can be exceedingly costly. Nevertheless, a trial-and-error approach can reliably identify opportunities for improving the look-ahead thread and quantify the benefits. A framework based on genetic algorithm can help search for the right set of changes to the look-ahead thread. In the set of applications where the speed of look-ahead has become the new limit, this method is found to improve the overall system performance by up to 1.48x with a geometric mean of 1.14x over the baseline decoupled look-ahead system, while reducing energy consumption by 11%.