SIMD defragmenter: efficient ILP realization on data-parallel architectures

SIMD defragmenter: efficient ILP realization on data-parallel architectures
复制标题

SIMD 碎片整理程序:数据并行架构上的高效 ILP 实现

DOI:
--
复制
发表时间:
2012
期刊:
ASPLOS XVII
影响因子:
--
通讯作者:
S. Mahlke
S. Mahlke
中科院分区:
--
文献类型:
--
作者:
Yongjun Park;Sangwon Seo;Hyunchul Park;Hyoun Kyu Cho;S. Mahlke

文献摘要

被引文献

相似文献

单个指令多数据(SIMD)加速器提供了一个节能的平台,以扩展移动系统的性能,同时仍保持后编程性。核心挑战是将SIMD硬件的并行资源转化为真实的应用程序性能。在科学应用中,自动矢量化技术已被证明在提取大量数据级并行性(DLP)方面非常有效。但是,由于较低的旅行计数循环,复杂的控制流和不均匀的执行行为,矢量化通常对媒体应用的有效性差得多。结果,由于DLP不足,SIMD车道保持空闲状态。为了攻击这个问题,本文提出了一个名为Simd Defragmenter的新矢量化通行证,以揭示隐藏的DLP,该DLP以指令级并行性(ILP)的形式潜伏在表面下方。困难是管理数据包装/开箱开销,这很容易超过SIMD执行所获得的好处。 SIMD DEGRAGMENTER通过识别可以在SIMD车道上并行执行的兼容指令(子图)组来克服此问题。通过在子图级别的批量化合物中,包装/解开包装开销就可以最小化。在16车道的SIMD处理器上,实验结果表明,SIMD碎片部在传统的环矢量中达到平均1.6倍的速度,并且比先前的研究方法获得了31%的增​​长,以将ILP转换为DLP。
Single-instruction multiple-data (SIMD) accelerators provide an energy-efficient platform to scale the performance of mobile systems while still retaining post-programmability. The central challenge is translating the parallel resources of the SIMD hardware into real application performance. In scientific applications, automatic vectorization techniques have proven quite effective at extracting large levels of data-level parallelism (DLP). However, vectorization is often much less effective for media applications due to low trip count loops, complex control flow, and non-uniform execution behavior. As a result, SIMD lanes remain idle due to insufficient DLP. To attack this problem, this paper proposes a new vectorization pass called SIMD Defragmenter to uncover hidden DLP that lurks below the surface in the form of instruction-level parallelism (ILP). The difficulty is managing the data packing/unpacking overhead that can easily exceed the benefits gained through SIMD execution. The SIMD degragmenter overcomes this problem by identifying groups of compatible instructions (subgraphs) that can be executed in parallel across the SIMD lanes. By SIMDizing in bulk at the subgraph level, packing/unpacking overhead is minimized. On a 16-lane SIMD processor, experimental results show that SIMD defragmentation achieves a mean 1.6x speedup over traditional loop vectorization and a 31% gain over prior research approaches for converting ILP to DLP.