Accelerating Sequential Applications on CMPs Using Core Spilling

Accelerating Sequential Applications on CMPs Using Core Spilling
复制标题

使用核心溢出加速 CMP 上的顺序应用

DOI:
10.1109/tpds.2007.1085
复制
发表时间:
2007
影响因子:
5.3
通讯作者:
Krzysztof Rutkowski
Krzysztof Rutkowski
中科院分区:
计算机科学2区
文献类型:
--
作者:
J. Cong;Guoling Han;Ashok Jagannathan;G. Reinman;Krzysztof Rutkowski

文献摘要

被引文献

相似文献

芯片多处理器 (CMP) 提供了一种为多任务或多线程应用程序利用线程级并行性的可扩展方法。然而,单线程应用程序将难以动态利用 CMP 中的静态分区资源。此类顺序应用程序可能难以静态分解为线程,或者可能只是无法重新编译或重新编译的遗留代码。我们提出了一种新颖的方法来动态加速多核上顺序应用程序的性能。当一个核心上的资源耗尽时,允许执行从一个核心溢出到另一个核心。我们提出了两种技术来实现内核之间的低开销迁移:预溢出和基于位置的过滤。我们开发并分析了一种仲裁机制,可以在 CMP 上的一组顺序应用程序之间智能地分配内核。平均而言,八核 CMP 上的核心溢出可以将单线程性能提高 35%。我们进一步探索运行由整个 SPEC 2000 基准套件以各种组合和到达时间组成的多个应用程序工作负载的八核 CMP。在有空闲核心的情况下,使用核心溢出来加速当前正在运行的应用程序集,我们可以将性能提高高达 40%。
Chip multiprocessors (CMPs) provide a scalable means of exploiting thread-level parallelism for multitasking or multithreaded applications. However, single-threaded applications will have difficulty dynamically leveraging the statically partitioned resources in a CMP. Such sequential applications may be difficult to statically decompose into threads or may simply be-a legacy code where recompilation is not possible or cost-effective. We present a novel approach to dynamically accelerate the performance of sequential application(s) on multiple cores. Execution is allowed to spill from one core to another when resources on one core have been exhausted. We propose two techniques to enable low-overhead migration between cores: prespilling and locality-based filtering. We develop and analyze an arbitration mechanism to intelligently allocate cores among a set of sequential applications on a CMP. On average, core spilling on an eight-core CMP can accelerate single-threaded performance by 35 percent. We further explore an eight-core CMP running a multiple application workload composed of the entire SPEC 2000 benchmark suite in various combinations and arrival times. Using core spilling to accelerate the current set of running applications in cases where there are idle cores, we achieve up to a 40 percent improvement in performance.