Lattice Boltzmann simulation optimization on leading multicore platforms

Lattice Boltzmann simulation optimization on leading multicore platforms
复制标题

领先多核平台上的莱迪思玻尔兹曼仿真优化

DOI:
--
复制
发表时间:
2008
期刊:
2008 IEEE International Symposium on Parallel and Distributed Processing
影响因子:
--
通讯作者:
K. Yelick
K. Yelick
中科院分区:
--
文献类型:
--
作者:
Samuel Williams;J. Carter;L. Oliker;J. Shalf;K. Yelick

文献摘要

被引文献

相似文献

我们提出了一种自动调优方法来优化新兴多核架构上的应用程序性能。该方法将基于搜索的性能优化思想(在线性代数和FFT库中很流行)扩展到特定于应用程序的计算内核。我们的工作将此策略应用于晶格玻尔兹曼应用程序(LBMHD),由于其复杂的数据结构和内存访问模式,该应用程序在历史上对标量微处理器的使用很差。我们探索了HPC文献中最广泛的多核架构之一,包括英特尔Clovertown, AMD Opteron X2, Sun Niagara!以及英特尔的itanium 2单核处理器。我们不是为每个系统手动调优LBMHD,而是开发一个代码生成器,它允许我们为每个平台确定高度优化的版本,同时分摊人工编程工作。结果表明,我们的自调优LBMHD应用程序比原始代码提高了14倍。此外,我们还详细分析了每种优化,揭示了未来多核系统和应用程序的惊人硬件瓶颈和软件挑战。
We present an auto-tuning approach to optimize application performance on emerging multicore architectures. The methodology extends the idea of search-based performance optimizations, popular in linear algebra and FFT libraries, to application-specific computational kernels. Our work applies this strategy to a lattice Boltzmann application (LBMHD) that historically has made poor use of scalar microprocessors due to its complex data structures and memory access patterns. We explore one of the broadest sets of multicore architectures in the HPC literature, including the Intel Clovertown, AMD Opteron X2, Sun Niagara!, STI Cell, as well as the single core Intel Itanium.2. Rather than hand-tuning LBMHD for each system, we develop a code generator that allows us identify a highly optimized version for each platform, while amortizing the human programming effort. Results show that our auto- tuned LBMHD application achieves up to a 14times improvement compared with the original code. Additionally, we present detailed analysis of each optimization, which reveal surprising hardware bottlenecks and software challenges for future multicore systems and applications.