Parallel grid library for rapid and flexible simulation development

Parallel grid library for rapid and flexible simulation development
复制标题

用于快速灵活仿真开发的并行网格库

DOI:
10.1016/j.cpc.2012.12.017
复制
发表时间:
2012
期刊:
ArXiv
影响因子:
--
通讯作者:
M. Palmroth
M. Palmroth
中科院分区:
--
文献类型:
--
作者:
I. Honkonen;S. Alfthan;A. Sandroos;P. Janhunen;M. Palmroth

文献摘要

被引文献

相似文献

我们提出了一个易于使用和灵活的网格库开发高度可扩展的并行仿真。分布式CarbonaceCell-Refinable Grid(dccrg)支持自适应网格细化,并允许使用任意C++类作为单元数据。网格单元中的数据量可以在空间和时间上变化,从而允许dccrg用于非常不同类型的模拟,例如流体和粒子代码。Dccrg在不同进程上的相邻单元之间透明地、异步地传输数据,允许计算和通信重叠。这使得出色的可扩展性至少高达32 k核心的磁流体动力学测试取决于问题和硬件。在这里介绍的dccrg版本中,部分网格元数据在MPI进程之间复制,将自适应网格细化(AMR)的可扩展性减少到200到600个进程之间。Dccrg是一个免费软件,任何人都可以使用,学习和修改,并在https://gitorious.org/dccrg上提供。用户也请引用这项工作时发表的结果与dccrg。项目摘要:项目名称:DCCRG目录标识符:AEOM_v1_0项目摘要URL:http://cpc.cs.qub.ac.uk/summaries/AEOM_v1_0.html项目可从:CPC项目图书馆,皇后大学,贝尔法斯特,N。爱尔兰许可条款:GNU Lesser General Public License version 3分布式程序的行数,包括测试数据等:54975分布式程序中的字节数,包括测试数据等:974015分发格式:tar.gz编程语言:C++.计算机:PC、集群、超级计算机。操作系统:POSIX。代码已使用MPI并行化,并使用1-32768个进程进行了测试RAM:每个进程10 MB-10 GB分类:4.12、4.14、6.5、19.3、19.10、20。外部例程:MPI-2 [1],boost [2],Zoltan [3],sfc++ [4]问题的性质:网格库支持网格单元中的任意数据,并行自适应网格细化,透明的远程邻居数据更新和负载平衡。解决方法:模拟网格由邻接列表(图)表示,其中顶点存储在哈希表中,边缘存储在连续数组中。消息传递接口标准用于并行化。单元格数据在实例化网格时作为模板参数给出。限制:逻辑上的网格。运行时间:运行时间取决于硬件,问题和解决方法。小问题可以在一分钟内解决,大问题可能需要几周时间。使用默认选项,包中提供的示例和测试只需要不到一分钟的时间。在这里介绍的dccrg版本中,自适应网格细化的速度最多为每秒创建106个总单元的数量级。参考文献:
We present an easy to use and flexible grid library for developing highly scalable parallel simulations. The distributed cartesian cell-refinable grid (dccrg) supports adaptive mesh refinement and allows an arbitrary C++ class to be used as cell data. The amount of data in grid cells can vary both in space and time allowing dccrg to be used in very different types of simulations, for example in fluid and particle codes. Dccrg transfers the data between neighboring cells on different processes transparently and asynchronously allowing one to overlap computation and communication. This enables excellent scalability at least up to 32 k cores in magnetohydrodynamic tests depending on the problem and hardware. In the version of dccrg presented here part of the mesh metadata is replicated between MPI processes reducing the scalability of adaptive mesh refinement (AMR) to between 200 and 600 processes. Dccrg is free software that anyone can use, study and modify and is available at https://gitorious.org/dccrg. Users are also kindly requested to cite this work when publishing results obtained with dccrg. PROGRAM SUMMARY: Program title: DCCRG Catalogue identifier: AEOM_v1_0 Program summary URL:http://cpc.cs.qub.ac.uk/summaries/AEOM_v1_0.html Program obtainable from: CPC Program Library, Queen’s University, Belfast, N. Ireland Licensing provisions: GNU Lesser General Public License version 3 No. of lines in distributed program, including test data, etc.: 54975 No. of bytes in distributed program, including test data, etc.: 974015 Distribution format: tar.gz Programming language: C++. Computer: PC, cluster, supercomputer. Operating system: POSIX. The code has been parallelized using MPI and tested with 1–32768 processes RAM: 10 MB–10 GB per process Classification: 4.12, 4.14, 6.5, 19.3, 19.10, 20. External routines: MPI-2 [1], boost [2], Zoltan [3], sfc++ [4] Nature of problem: Grid library supporting arbitrary data in grid cells, parallel adaptive mesh refinement, transparent remote neighbor data updates and load balancing. Solution method: The simulation grid is represented by an adjacency list (graph) with vertices stored into a hash table and edges into contiguous arrays. Message Passing Interface standard is used for parallelization. Cell data is given as a template parameter when instantiating the grid. Restrictions: Logically cartesian grid. Running time: Running time depends on the hardware, problem and the solution method. Small problems can be solved in under a minute and very large problems can take weeks. The examples and tests provided with the package take less than about one minute using default options. In the version of dccrg presented here the speed of adaptive mesh refinement is at most of the order of 106total created cells per second. References: