Characterizing the performance of intel optane persistent memory: a close look at its on-DIMM buffering

Characterizing the performance of intel optane persistent memory: a close look at its on-DIMM buffering
复制标题

DOI:
10.1145/3492321.3519556
复制
发表时间:
2022-03
期刊:
Proceedings of the Seventeenth European Conference on Computer Systems
影响因子:
--
通讯作者:
Lingfeng Xiang;Xingsheng Zhao;J. Rao;Song Jiang;Hong Jiang
Lingfeng Xiang;Xingsheng Zhao;J. Rao;Song Jiang;Hong Jiang
中科院分区:
其他
文献类型:
--
作者:
Lingfeng Xiang;Xingsheng Zhao;J. Rao;Song Jiang;Hong Jiang

文献摘要

相似文献

我们提出了一项对英特尔Optane DC持续记忆(DCPMM)的全面而深入的研究。我们的重点是探索Optane的dimm读取 - 写缓冲的内部设计及其对应用程序感知性能,读写放大的影响,不同类型的持久性的开销以及持久性模型之间的权衡。尽管我们的测量结果证实了现有的分析研究的结果,但我们有了新的发现并提供了新的见解。值得注意的是,我们发现在单独的on-dimm读写缓冲区中,读取和写入的管理方式有所不同。这两个缓冲区的尺寸可比,具有不同的目的。读取缓冲液具有更高的并发性和有效的在DIMM预取,可导致较高的读带宽和出色的顺序性能。但是,它无助于隐藏媒体访问延迟。相比之下,写缓冲区提供有限的并发性,但在管道中是支持异步写入DDR-T协议中的关键阶段。出乎意料的是,除了写入合并外,写缓冲区的提供比读取还低,并且不管工作集的大小,写入类型,访问模式或持久性模型,而写入延迟。此外,我们发现,Cacheline访问粒度与3D-XPOINT媒体访问粒度之间的不匹配对CPU缓存预取的有效性产生了负面影响,并导致浪费持续的内存带宽。我们的主张是在持续程序的绩效分析和优化中解除读写。我们基于这种见解提供了三个案例研究,并证明了大量绩效的改进。我们在两代Optane DCPMM上验证结果。
We present a comprehensive and in-depth study of Intel Optane DC persistent memory (DCPMM). Our focus is on exploring the internal design of Optane's on-DIMM read-write buffering and its impacts on application-perceived performance, read and write amplifications, the overhead of different types of persists, and the tradeoffs between persistency models. While our measurements confirm the results of the existing profiling studies, we have new discoveries and offer new insights. Notably, we find that read and write are managed differently in separate on-DIMM read and write buffers. Comparable in size, the two buffers serve distinct purposes. The read buffer offers higher concurrency and effective on-DIMM prefetching, leading to high read bandwidth and superior sequential performance. However, it does not help hide media access latency. In contrast, the write buffer offers limited concurrency but is a critical stage in a pipeline that supports asynchronous write in the DDR-T protocol. Surprisingly, in addition to write coalescing, the write buffer delivers lower than read and consistent write latency regardless of the working set size, the type of write, the access pattern, or the persistency model. Furthermore, we discover that the mismatch between cacheline access granularity and the 3D-Xpoint media access granularity negatively impacts the effectiveness of CPU cache prefetching and leads to wasted persistent memory bandwidth. Our proposition is to decouple read and write in the performance analysis and optimization of persistent programs. We present three case studies based on this insight and demonstrate considerable performance improvements. We verify the results on two generations of Optane DCPMM.