RDMA vs. RPC for Implementing Distributed Data Structures

RDMA vs. RPC for Implementing Distributed Data Structures
复制标题

DOI:
10.1109/ia349570.2019.00009
复制
发表时间:
2019-10
期刊:
2019 IEEE/ACM 9th Workshop on Irregular Applications: Architectures and Algorithms (IA3)
影响因子:
--
通讯作者:
Benjamin Brock;Yuxin Chen;Jiakun Yan;John Douglas Owens;A. Buluç;K. Yelick
Benjamin Brock;Yuxin Chen;Jiakun Yan;John Douglas Owens;A. Buluç;K. Yelick
中科院分区:
其他
文献类型:
--
作者:
Benjamin Brock;Yuxin Chen;Jiakun Yan;John Douglas Owens;A. Buluç;K. Yelick

文献摘要

相似文献

分布式数据结构是实现科学模拟和数据分析的可扩展应用程序的关键。在本文中,我们研究分布式数据结构的两种实现方式:远程直接内存访问(RDMA)和远程过程调用(RPC)。我们关注需要单独访问分布式数据结构的远程部分的操作,例如访问哈希表桶或分布式队列,而不是所有处理器共同交换信息的全局操作。我们通过微基准测试和近似每种风格成本的性能模型来研究两种风格之间的权衡。 RDMA 操作在网络中具有直接硬件支持,因此延迟和开销较低,而 RPC 操作更具表现力,但成本较高,并且可能会受到远程端缺乏关注的影响。我们还进行了实验,将基于 RDMA 和 RPC 的数据结构操作的实际性能与预测性能进行比较,以评估我们模型的准确性,并表明,虽然该模型并不总是精确预测运行时间,但它允许我们在所示示例中选择最佳实现。我们相信这种分析将帮助开发人员设计在当前网络架构上表现良好的数据结构,并帮助网络架构师为此类分布式数据结构提供更好的支持。
Distributed data structures are key to implementing scalable applications for scientific simulations and data analysis. In this paper we look at two implementation styles for distributed data structures: remote direct memory access (RDMA) and remote procedure call (RPC). We focus on operations that require individual accesses to remote portions of a distributed data structure, e.g., accessing a hash table bucket or distributed queue, rather than global operations in which all processors collectively exchange information. We look at the trade-offs between the two styles through microbenchmarks and a performance model that approximates the cost of each. The RDMA operations have direct hardware support in the network and therefore lower latency and overhead, while the RPC operations are more expressive but higher cost and can suffer from lack of attentiveness from the remote side. We also run experiments to compare the real-world performance of RDMA- and RPC-based data structure operations with the predicted performance to evaluate the accuracy of our model, and show that while the model does not always precisely predict running time, it allows us to choose the best implementation in the examples shown. We believe this analysis will assist developers in designing data structures that will perform well on current network architectures, as well as network architects in providing better support for this class of distributed data structures.