The potential for using thread-level data speculation to facilitate automatic parallelization

The potential for using thread-level data speculation to facilitate automatic parallelization
复制标题

DOI:
10.1109/hpca.1998.650541
复制
发表时间:
1998-01
期刊:
Proceedings 1998 Fourth International Symposium on High-Performance Computer Architecture
影响因子:
--
通讯作者:
J. Steffan;T. Mowry
J. Steffan;T. Mowry
中科院分区:
其他
文献类型:
--
作者:
J. Steffan;T. Mowry

文献摘要

被引文献

相似文献

当我们展望未来,以及一个芯片上有十亿个晶体管的前景时,微处理器似乎不可避免地将利用具有多个并行线程。然而,为了实现这些“单芯片多处理器”的全部潜力,我们必须找到一种方法来并行化非数值应用程序。不幸的是,编译器在并行化非数字代码方面几乎没有成功,因为它们的访问模式很复杂。本文探讨了潜在的线程级数据推测(TLDS),以克服这一限制,允许编译器仅将并行化视为成本/效益的权衡,而不是一些可能违反程序的正确性。我们的实验结果表明,与现实的编译器支持,TLDS可以提供显着的程序加速。我们还表明,通过适度的硬件扩展,一个通用的单芯片多处理器可以支持TLDS通过增强其高速缓存一致性方案来检测依赖违规,并通过使用主数据缓存缓冲投机状态。
As we look to the future, and the prospect of a billion transistors on a chip, it seems inevitable that microprocessors will exploit having multiple parallel threads. To achieve the full potential of these "single-chip multiprocessors", however, we must find a way to parallelize non-numeric applications. Unfortunately, compilers have had little success in parallelizing non-numeric codes due to their complex access patterns. This paper explores the potential for using thread-level data speculation (TLDS) to overcome this limitation by allowing the compiler to view parallelization solely as a cost/benefit tradeoff rather than something which is likely to violate program correctness. Our experimental results demonstrate that with realistic compiler support, TLDS can offer significant program speedups. We also demonstrate that through modest hardware extensions, a generic single-chip multiprocessor could support TLDS by augmenting its cache coherence scheme to detect dependence violations, and by using the primary data caches to buffer speculative state.