A unified model for multicore architectures

A unified model for multicore architectures
复制标题

多核架构的统一模型

DOI:
10.1145/1463768.1463780
复制
发表时间:
2008
期刊:
J. Parallel Distributed Comput.
影响因子:
--
通讯作者:
M. Zubair
M. Zubair
中科院分区:
--
文献类型:
--
作者:
J. Savage;M. Zubair

文献摘要

被引文献

相似文献

随着多核心和许多核心体系结构的出现,我们面临着一个对并行计算的新问题,即层次并行caches的管理。所有早期模型的一个主要局限性是它们无法建模不同级别的库中共享程度不同的多层处理器。我们提出了一个统一的内存层次结构模型,该模型解决了这些局限性,并且是针对具有多内存层次结构的单个处理器开发的MHG模型的扩展。我们证明,我们的统一框架可以应用于多种应用程序的多项多层架构。特别是,我们在层次结构中的不同级别之间的内存流量得出了下限,以进行财务和科学计算。我们还为财务应用程序提供了多功能算法,该算法在不同的高速缓存级别之间表现出恒定的因素最佳内存流量。我们在具有两个四核Intel Xeon 5310 1.6GHz处理器的多核心系统上实现了算法,总共具有8个核心。我们的算法优于编译器优化和自动合行的代码,其倍数高达7.3。
With the advent of multicore and many core architectures, we are facing a problem that is new to parallel computing, namely, the management of hierarchical parallel caches. One major limitation of all earlier models is their inability to model multicore processors with varying degrees of sharing of caches at different levels. We propose a unified memory hierarchy model that addresses these limitations and is an extension of the MHG model developed for a single processor with multi-memory hierarchy. We demonstrate that our unified framework can be applied to a number of multicore architectures for a variety of applications. In particular, we derive lower bounds on memory traffic between different levels in the hierarchy for financial and scientific computations. We also give a multicore algorithms for a financial application that exhibits a constant-factor optimal amount of memory traffic between different cache levels. We implemented the algorithm on a multicore system with two Quad-Core Intel Xeon 5310 1.6GHz processors having a total of 8 cores. Our algorithms outperform compiler optimized and auto-parallelized code by a factor of up to 7.3.