Persistence Parallelism Optimization: A Holistic Approach from Memory Bus to RDMA Network

Persistence Parallelism Optimization: A Holistic Approach from Memory Bus to RDMA Network
复制标题

DOI:
10.1109/micro.2018.00047
复制
发表时间:
2018-10
期刊:
2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Xing Hu;Matheus A. Ogleari;Jishen Zhao;Shuangchen Li;Abanti Basak;Yuan Xie
Xing Hu;Matheus A. Ogleari;Jishen Zhao;Shuangchen Li;Abanti Basak;Yuan Xie
中科院分区:
其他
文献类型:
--
作者:
Xing Hu;Matheus A. Ogleari;Jishen Zhao;Shuangchen Li;Abanti Basak;Yuan Xie

文献摘要

被引文献

相似文献

新兴的非易失性存储器(NVM),诸如相变存储器(PCM)和电阻式RAM(ReRAM),结合了快速字节寻址能力和数据持久性的特征,这对于诸如文件系统和数据库的数据服务是有益的。为了支持数据持久性,持久存储器系统需要对写入请求进行排序。持久请求的数据路径由三段组成:通过该高速缓存层次结构到内存控制器,通过总线从内存控制器到内存设备,以及通过网络从远程节点到本地节点。以前的工作有助于显着提高持久性并行在第一段的数据path.However,我们观察到,存储器总线和远程直接存储器访问(RDMA)网络仍然严重利用不足,因为在这两个段的持久性并行在排序过程中没有充分利用。在本文中,我们提出了一种新的架构,以进一步提高持久性并行存储器总线和RDMA网络。首先,我们利用线程间的持久性并行屏障历元管理更好的银行级并行(BLP)。第二,我们启用线程内持久化并行远程请求通过RDMA网络与缓冲严格持久化。有了这些特性,该架构可以有效地支持写入数据路径所有三个部分的持久性。实验结果表明,对于本地应用,与原有的缓冲持久化工作相比,该机制可以获得1.3倍的性能提升.此外,它还可以为通过RDMA网络提供服务的远程应用程序实现1.93倍的性能提升。
Emerging non-volatile memories (NVM), such as phase change memory (PCM) and Resistive RAM (ReRAM), incorporate the features of fast byte-addressability and data persistence, which are beneficial for data services such as file systems and databases. To support data persistence, a persistent memory system requires ordering for write requests. The datapath of a persistent request consists of three segments: through the cache hierarchy to the memory controller, through the bus from the memory controller to memory devices, and through the network from a remote node to a local node. Previous work contributes significantly to improve the persistence parallelism in the first segment of the data path. However, we observe that the memory bus and the Remote Direct Memory Access (RDMA) network remain severely under-utilized because the persistence parallelism in these two segments is not fully leveraged during ordering. In this paper, we propose a novel architecture to further improve the persistence parallelism in the memory bus and the RDMA network. First, we utilize inter-thread persistence parallelism for barrier epoch management with better bank-level parallelism (BLP). Second, we enable intra-thread persistence parallelism for remote requests through RDMA network with buffered strict persistence. With these features, the architecture efficiently supports persistence through all three segments of the write datapath. Experimental results show that for local applications, the proposed mechanism can achieve 1.3x performance improvement, compared to the original buffered persistence work. In addition, it can achieve 1.93x performance improvement for remote applications serviced through the RDMA network.