Process-in-process: techniques for practical address-space sharing

Process-in-process: techniques for practical address-space sharing
复制标题

进程中进程:实用地址空间共享技术

DOI:
--
复制
发表时间:
2018
期刊:
IEEE International Symposium on High-Performance Parallel Distributed Computing
影响因子:
--
通讯作者:
Y. Ishikawa
Y. Ishikawa
中科院分区:
--
文献类型:
--
作者:
A. Hori;Min Si;Balazs Gerofi;Masamichi Takagi;Jai Dayal;P. Balaji;Y. Ishikawa

文献摘要

被引文献

相似文献

当今众核CPU的两种最常见的并行执行模型是多进程(例如,MPI)和多线程(例如,OpenMP)。多进程模型允许每个进程拥有一个私有地址空间,尽管进程可以显式地分配共享内存区域。多线程模型默认情况下共享所有地址空间,尽管线程可以显式地将数据移动到线程专用存储。在本文中,我们提出了第三个模型称为进程中进程(PiP),其中多个进程被映射到一个虚拟地址空间。因此,每个进程仍然拥有其进程私有存储(如多进程模型),但可以直接访问同一虚拟地址空间中其他进程的私有存储(如多线程模型)。多个进程之间共享地址空间的想法本身并不新鲜。然而,PiP的独特之处在于它的设计完全在用户空间中,使其成为大型超级计算系统的便携式和实用方法,其中移植现有的基于操作系统的技术可能很难。PiP库结构紧凑,旨在与其他运行时系统(如MPI和OpenMP)集成,作为提高HPC应用程序通信性能的便携式低级支持。我们通过各种并行运行时优化和在数据分析应用程序中的直接使用来展示PiP环境的独特性。我们评估了几个平台上的PiP,包括两个高级别的超级计算机,我们测量和分析的PiP的性能,通过使用各种微和宏内核,代理应用程序以及数据分析应用程序。
The two most common parallel execution models for many-core CPUs today are multiprocess (e.g., MPI) and multithread (e.g., OpenMP). The multiprocess model allows each process to own a private address space, although processes can explicitly allocate shared-memory regions. The multithreaded model shares all address space by default, although threads can explicitly move data to thread-private storage. In this paper, we present a third model called process-in-process (PiP), where multiple processes are mapped into a single virtual address space. Thus, each process still owns its process-private storage (like the multiprocess model) but can directly access the private storage of other processes in the same virtual address space (like the multithread model). The idea of address-space sharing between multiple processes itself is not new. What makes PiP unique, however, is that its design is completely in user space, making it a portable and practical approach for large supercomputing systems where porting existing OS-based techniques might be hard. The PiP library is compact and is designed for integrating with other runtime systems such as MPI and OpenMP as a portable low-level support for boosting communication performance in HPC applications. We showcase the uniqueness of the PiP environment through both a variety of parallel runtime optimizations and direct use in a data analysis application. We evaluate PiP on several platforms including two high-ranking supercomputers, and we measure and analyze the performance of PiP by using a variety of micro- and macro-kernels, a proxy application as well as a data analysis application.