MPI+Threads: runtime contention and remedies

MPI+Threads: runtime contention and remedies
复制标题

DOI:
10.1145/2688500.2688522
复制
发表时间:
2015-01
期刊:
Proceedings of the 20th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming
影响因子:
--
通讯作者:
A. Amer;Huiwei Lu;Yanjie Wei;P. Balaji;S. Matsuoka
A. Amer;Huiwei Lu;Yanjie Wei;P. Balaji;S. Matsuoka
中科院分区:
其他
文献类型:
--
作者:
A. Amer;Huiwei Lu;Yanjie Wei;P. Balaji;S. Matsuoka

文献摘要

相似文献

混合MPI+线程编程已经成为“MPI无处不在”模型的替代模型,以更好地处理集群节点中不断增加的核心密度。虽然MPI标准允许多线程并发通信,但这种灵活性伴随着在MPI实现中维护线程安全的成本,通常使用临界区实现。在以前的作品,研究了MPI实现中的关键部分粒度的重要性,在本文中,我们调查的影响,关键部分仲裁通信性能。我们首先分析MPI运行时多线程并发通信发生在分层内存系统。我们的研究结果表明,大多数MPI实现使用的基于互斥的方法,今天可能会招致性能损失,由于不公平的仲裁。然后,我们提出的方法,以减轻这些处罚,先到先得的仲裁和优先级锁定方案,有利于线程做有用的工作。通过使用几个基准测试和应用程序的评估,我们展示了高达5倍的性能改进。
Hybrid MPI+Threads programming has emerged as an alternative model to the “MPI everywhere” model to better handle the increasing core density in cluster nodes. While the MPI standard allows multithreaded concurrent communication, such flexibility comes with the cost of maintaining thread safety within the MPI implementation, typically implemented using critical sections. In contrast to previous works that studied the importance of critical-section granularity in MPI implementations, in this paper we investigate the implication of critical-section arbitration on communication performance. We first analyze the MPI runtime when multithreaded concurrent communication takes place on hierarchical memory systems. Our results indicate that the mutex-based approach that most MPI implementations use today can incur performance penalties due to unfair arbitration. We then present methods to mitigate these penalties with a first-come, first-served arbitration and a priority locking scheme that favors threads doing useful work. Through evaluations using several benchmarks and applications, we demonstrate up to 5-fold improvement in performance.