Studies on the energy and deep memory behaviour of a cache-oblivious, task-based hyperbolic PDE solver

Studies on the energy and deep memory behaviour of a cache-oblivious, task-based hyperbolic PDE solver
复制标题

研究缓存无关、基于任务的双曲 PDE 求解器的能量和深度内存行为

DOI:
10.1177/1094342019842645
复制
发表时间:
2018
期刊:
The International Journal of High Performance Computing Applications
影响因子:
--
通讯作者:
T. Weinzierl
T. Weinzierl
中科院分区:
--
文献类型:
--
作者:
D. E. Charrier;B. Hazelwood;Ekaterina Tutlyaeva;M. Bader;M. Dumbser;A. Kudryavtsev;A. Moskovsky;T. Weinzierl

文献摘要

被引文献

相似文献

我们研究了使用ExaHyPE引擎的地震模拟的性能行为,特别关注内存特征和能量需求。ExaHyPE将动态自适应网格细化(AMR)与ADER-DG相结合。它使用任务进行并行化,并且高速缓存效率很高。AMR加上ADER-DG产生了一个本质上高度动态的任务图,它既包括算术代价高昂的任务,也包括挑战内存延迟的任务。昂贵的任务和整个代码受益于AVX矢量化,尽管我们遭受了内存访问猝发。芯片频率的降低提高了代码的能量比。然而,它并不能缓解突发效应。一旦我们添加了英特尔Optane技术,显著增加了核心计数,或者使单独的、计算繁重的任务从封闭缓存中消失,猝发的延迟损失就会变得更糟。线程超额预定来隐藏这些延迟惩罚对于非包含性缓存来说会适得其反,因为它破坏了缓存和矢量化特性。在内存密集型和计算代价高昂的任务重叠的情况下,ExaHyPE的缓存无关实现仍然可以有效地利用深度、非包含性、异类内存,因为主内存未命中很少出现,并且只会减慢几个核心的速度。因此,我们提出了具有动态、非同构任务图的超级计算模拟代码在混合不同计算特征的任务时得到线程运行时的主动支持,并且我们提出未来的硬件主动允许代码降低运行特定任务类型的核心的频率。
We study the performance behaviour of a seismic simulation using the ExaHyPE engine with a specific focus on memory characteristics and energy needs. ExaHyPE combines dynamically adaptive mesh refinement (AMR) with ADER-DG. It is parallelized using tasks, and it is cache efficient. AMR plus ADER-DG yields a task graph which is highly dynamic in nature and comprises both arithmetically expensive tasks and tasks which challenge the memory’s latency. The expensive tasks and thus the whole code benefit from AVX vectorization, although we suffer from memory access bursts. A frequency reduction of the chip improves the code’s energy-to-solution. Yet, it does not mitigate burst effects. The bursts’ latency penalty becomes worse once we add Intel Optane technology, increase the core count significantly or make individual, computationally heavy tasks fall out of close caches. Thread overbooking to hide away these latency penalties becomes contra-productive with noninclusive caches as it destroys the cache and vectorization character. In cases where memory-intense and computationally expensive tasks overlap, ExaHyPE’s cache-oblivious implementation nevertheless can exploit deep, noninclusive, heterogeneous memory effectively, as main memory misses arise infrequently and slow down only few cores. We thus propose that upcoming supercomputing simulation codes with dynamic, inhomogeneous task graphs are actively supported by thread runtimes in intermixing tasks of different compute character, and we propose that future hardware actively allows codes to downclock the cores running particular task types.