SoAx: A generic C++ Structure of Arrays for handling particles in HPC codes

SoAx: A generic C++ Structure of Arrays for handling particles in HPC codes
复制标题

DOI:
10.1016/j.cpc.2017.11.015
复制
发表时间:
2017-10
期刊:
ArXiv
影响因子:
--
通讯作者:
H. Homann;Francois Laenen
H. Homann;Francois Laenen
中科院分区:
其他
文献类型:
--
作者:
H. Homann;Francois Laenen

文献摘要

被引文献

相似文献

物理问题的数值研究常常需要根据一组给定的方程对大量粒子的动力学进行积分。粒子的特征是它们所携带的信息,如身份、位置等。一般来说,在高性能计算(HPC)代码中处理粒子有两种不同的可能性。结构数组(AoS)的概念符合面向对象编程(OOP)范式的精神,因为粒子信息被实现为结构。在这里,一个对象(结构的实现)代表一个粒子,许多粒子的集合存储在一个数组中。相反,使用阵列结构(SoA)的概念,单个结构包含多个数组,每个数组代表整个粒子集的一个属性(如身份)。由于AoS方法的方便性和灵活性,它通常在HPC代码中实现。然而,对于一类问题,我们知道SoA的性能要比aop好得多。我们在粒子问题中证实了这一观察结果。通过基准测试,我们发现在现代英特尔至强处理器上,SoA实现通常比AoS实现快几倍。在英特尔的MIC协处理器上,性能差距甚至达到了十倍。GPU计算也是如此,使用计算型和多用途GPU。结合性能和方便性,我们提出了具有最佳性能(在cpu、mic和gpu上)的库SoAx,同时提供与aop相同的方便性。为此,SoAx使用了现代c++设计技术,例如模板元编程,它允许为用户定义的异构数据结构自动生成代码。程序摘要程序标题:SoAxProgram Files doi:http://dx.doi.org/10.17632/m463pc4mv8.1Licensing provisions: gplv3编程语言:c++问题性质:数组结构(SoA)通常比结构数组(AoS)更快,而AoS更方便。这个库(SoAx)结合了两者的优点。通过c++(11)元模板编程,SoAx在提供非常方便的用户界面(包括面向对象的元素处理)和灵活性的同时实现了最大的性能(有效地使用向量单元和现代cpu的缓存)。它被设计用于在高性能数值模拟环境中处理类似列表的粒子集(类似于struct int id; double[3] pos; float[3] vel;;)。它可以应用于许多其他问题。解决方法:模板元编程、表达式模板
The numerical study of physical problems often require integrating the dynamics of a large number of particles evolving according to a given set of equations. Particles are characterized by the information they are carrying such as an identity, a position other. There are generally speaking two different possibilities for handling particles in high performance computing (HPC) codes. The concept of anArray of Structures(AoS) is in the spirit of the object-oriented programming (OOP) paradigm in that the particle information is implemented as a structure. Here, an object (realization of the structure) represents one particle and a set of many particles is stored in an array. In contrast, using the concept of aStructure of Arrays(SoA), a single structure holds several arrays each representing one property (such as the identity) of the whole set of particles.The AoS approach is often implemented in HPC codes due to its handiness and flexibility. For a class of problems, however, it is known that the performance of SoA is much better than that of AoS. We confirm this observation for our particle problem. Using a benchmark we show that on modern Intel Xeon processors the SoA implementation is typically several times faster than the AoS one. On Intel’s MIC co-processors the performance gap even attains a factor of ten. The same is true for GPU computing, using both computational and multi-purpose GPUs.Combining performance and handiness, we present the library SoAx that has optimal performance (on CPUs, MICs, and GPUs) while providing the same handiness as AoS. For this, SoAx uses modern C++ design techniques such template meta programming that allows to automatically generate code for user defined heterogeneous data structures.Program summaryProgram Title:SoAxProgram Files doi:http://dx.doi.org/10.17632/m463pc4mv8.1Licensing provisions:GPLv3Programming language:C++Nature of problem:Structures of arrays (SoA) are generally faster than arrays of structures (AoS) while AoS are more handy. This library (SoAx) combines the advantages of both. By means of C++(11) meta-template programming SoAx achieves maximal performance (efficient use of vector units and cache of modern CPUs) while providing a very convenient user interface (including object-oriented element handling) and flexibility. It has been designed to handle list-like sets of particles (similar to struct int id; double[3] pos; float[3] vel;;) in the context of high-performance numerical simulations. It can be applied to many other problems.Solution method:Template Metaprogramming, Expression Templates