High efficient parallel numerical surface wave model based on an irregular quasi-rectangular domain decomposition scheme

High efficient parallel numerical surface wave model based on an irregular quasi-rectangular domain decomposition scheme
复制标题

DOI:
10.1007/s11430-014-4842-3
复制
发表时间:
2014-06
期刊:
Science China Earth Sciences
影响因子:
--
通讯作者:
Wei Zhao;Zhenya Song;F. Qiao;Xunqiang Yin
Wei Zhao;Zhenya Song;F. Qiao;Xunqiang Yin
中科院分区:
其他
文献类型:
--
作者:
Wei Zhao;Zhenya Song;F. Qiao;Xunqiang Yin

文献摘要

被引文献

相似文献

为了实现全局MASNUM面波模型的高并行效率,基于消息传递接口(MPI)环境,开发并实施了不规则准矩形域分解算法以及相关的计算点序列化和数据交换方案。新的并行版本的表面波模型在国家超级计算济南中心的神威蓝光超级计算机平台上进行了并行计算测试。测试涉及四种水平分辨率,分别为1°×1°、(1/2)°×(1/2)°、(1/4)°×(1/4)°和(1/8)°×(1/8)°。这些测试都是在没有数据输入/输出(IO)的情况下进行的,测试中使用的处理器最大数量达到131072个。测试结果表明,不同分辨率的模型的计算速度均随着处理器数量的增加而增加。当处理器数量为基础处理器数量的4倍时,所有分辨率的并行效率均大于80%。当处理器数量为基础处理器数量的8倍时,分辨率为1°×1°、(1/2)°×(1/2)°、(1/4)°×(1/4)°测试的并行效率均大于80%,其中使用131072个处理器进行(1/8)°×(1/8)°分辨率测试的并行效率为62%,几乎是神威蓝光全部处理器。当处理器数量为基础处理器数量的24倍时,分辨率为1°×1°、(1/2)°×(1/2)°、(1/4)°×(1/4)°测试的并行效率分别为72%、62%和38%。加速比和并行效率表明,不规则的准矩形域分解和序列化方案为全局数值波浪模型带来了高并行效率和良好的可扩展性。
To achieve high parallel efficiency for the global MASNUM surface wave model, the algorithm of an irregular quasi-rectangular domain decomposition and related serializing of calculating points and data exchanging schemes are developed and conducted, based on the environment of Message Passing Interface (MPI). The new parallel version of the surface wave model is tested for parallel computing on the platform of the Sunway BlueLight supercomputer in the National Supercomputing Center in Jinan. The testing involves four horizontal resolutions, which are 1°×1°, (1/2)°×(1/2)°, (1/4)°×(1/4)°, and (1/8)°×(1/8)°. These tests are performed without data Input/Output (IO) and the maximum amount of processors used in these tests reaches to 131072. The testing results show that the computing speeds of the model with different resolutions are all increased with the increasing of numbers of processors. When the number of processors is four times that of the base processor number, the parallel efficiencies of all resolutions are greater than 80%. When the number of processors is eight times that of the base processor number, the parallel efficiency of tests with resolutions of 1°×1°, (1/2)°×(1/2)° and (1/4)°×(1/4)° is greater than 80%, and it is 62% for the test with a resolution of (1/8)°×(1/8)° using 131072 processors, which is the nearly all processors of Sunway BlueLight. When the processor’s number is 24 times that of the base processor number, the parallel efficiencies for tests with resolutions of 1°×1°, (1/2)°×(1/2)°, and (1/4)° ×(1/4)° are 72%, 62%, and 38%, respectively. The speedup and parallel efficiency indicate that the irregular quasi-rectangular domain decomposition and serialization schemes lead to high parallel efficiency and good scalability for a global numerical wave model.