Lattice QCD with Domain Decomposition on Intel® Xeon Phi Co-Processors

Lattice QCD with Domain Decomposition on Intel® Xeon Phi Co-Processors
复制标题

DOI:
10.1109/sc.2014.11
复制
发表时间:
2014-11
期刊:
SC14: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Simon Heybrock;B. Joó;Dhiraj D. Kalamkar;M. Smelyanskiy;K. Vaidyanathan;T. Wettig;P. Dubey
Simon Heybrock;B. Joó;Dhiraj D. Kalamkar;M. Smelyanskiy;K. Vaidyanathan;T. Wettig;P. Dubey
中科院分区:
其他
文献类型:
--
作者:
Simon Heybrock;B. Joó;Dhiraj D. Kalamkar;M. Smelyanskiy;K. Vaidyanathan;T. Wettig;P. Dubey

文献摘要

被引文献

相似文献

移动数据的成本与计算成本之间的差距不断增长,这使得在极端架构上设计迭代求解器的差异可以通过减少数据移动的替代算法来减轻此问题在Intel®XeonPhi Phi Phocersoner(KNC)群集上,在晶格量子式铬化物动力学和基于域分解的替代求解算法的情况下片上的片段缩放到knc的所有60个核心。标准求解器[1],我们的完整多节点域分解求解器强尺度对更多节点,并减少5倍的时间。
The gap between the cost of moving data and the cost of computing continues to grow, making it ever harder to design iterative solvers on extreme-scale architectures. This problem can be alleviated by alternative algorithms that reduce the amount of data movement. We investigate this in the context of Lattice Quantum Chromo dynamics and implement such an alternative solver algorithm, based on domain decomposition, on Intel® Xeon Phi co-processor (KNC) clusters. We demonstrate close-to-linear on-chip scaling to all 60 cores of the KNC. With a mix of single- and half-precision the domain-decomposition method sustains 400-500 Gflop/s per chip. Compared to an optimized KNC implementation of a standard solver [1], our full multi-node domain-decomposition solver strong-scales to more nodes and reduces the time-to-solution by a factor of 5.