An Analytical Study of Loop Tiling for a Large-Scale Unstructured Mesh Application

An Analytical Study of Loop Tiling for a Large-Scale Unstructured Mesh Application
复制标题

DOI:
10.1109/sc.companion.2012.68
复制
发表时间:
2012-11
期刊:
2012 SC Companion: High Performance Computing, Networking Storage and Analysis
影响因子:
--
通讯作者:
M. Giles;G. Mudalige;C. Bertolli;P. Kelly;E. László;I. Reguly
M. Giles;G. Mudalige;C. Bertolli;P. Kelly;E. László;I. Reguly
中科院分区:
其他
文献类型:
--
作者:
M. Giles;G. Mudalige;C. Bertolli;P. Kelly;E. László;I. Reguly

文献摘要

被引文献

相似文献

限制新兴多核和众核处理器性能的主要瓶颈越来越多地是数据在其不同核和主存之间的移动。随着内核数量的增加,越来越多的数据需要与内存交换,以保持它们的充分利用。这个关键瓶颈已经限制了处理器的实用性,以及我们利用增加的并行性来实现更高性能的能力。另一方面,大量的计算机科学研究存在于平铺技术(也称为稀疏平铺)上,用于减少数据传输。这些工作展示了如何避免日益增长的内存瓶颈,但困难在于将这些想法扩展到实际应用中。这些算法很快变得非常复杂,编译器很难自动检测机会并执行执行策略。专注于非结构化网格应用类,在本文中,我们提出了一个初步的分析调查平铺(或循环阻塞)算法在现实世界的工业CFD应用的性能优势。我们分析估计在此应用程序中的主要并行循环的通信或内存访问的减少,并定量预测现代多核和许多核心硬件上可以获得的性能优势。分析表明,在一般情况下,四个减少数据移动的因素可以通过平铺并行循环来实现。节省的主要部分来自临时或瞬态数据阵列的收缩,这些数据阵列不需要写回主存,而是将它们保存在现代处理器的最后一级缓存(LLC)中。
Increasingly, the main bottleneck limiting performance on emerging multi-core and many-core processors is the movement of data between its different cores and main memory. As the number of cores increases, more and more data needs to be exchanged with memory to keep them fully utilized. This critical bottleneck is already limiting the utility of processors and our ability to leverage increased parallelism to achieve higher performance. On the other hand, considerable computer science research exists on tiling techniques (also known as sparse tiling), for reducing data transfers. Such work demonstrates how the increasing memory bottleneck could be avoided but the difficulty has been in extending these ideas to real-world applications. These algorithms quickly become highly complicated, and it has be very difficult to for a compiler to automatically detect the opportunities and implement the execution strategy. Focusing on the unstructured mesh application class, in this paper, we present a preliminary analytical investigation into the performance benefits of tiling (or loop-blocking) algorithms on a realworld industrial CFD application. We analytically estimate the reductions in communications or memory accesses for the main parallel loops in this application and predict quantitatively the performance benefits that can be gained on modern multi-core and many core hardware. The analysis demonstrates that in general a factor of four reduction in data movement can be achieved by tiling parallel loops. A major part of the savings come from contraction of temporary or transient data arrays that need not be written back to main memory, by holding them in the last level cache (LLC) of modern processors.