HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVM

HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVM
复制标题

DOI:
10.1145/3477132.3483550
复制
发表时间:
2021-10
期刊:
Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles
影响因子:
--
通讯作者:
Amanda Raybuck;Tim Stamler;Wei Zhang;M. Erez;Simon Peter
Amanda Raybuck;Tim Stamler;Wei Zhang;M. Erez;Simon Peter
中科院分区:
其他
文献类型:
--
作者:
Amanda Raybuck;Tim Stamler;Wei Zhang;M. Erez;Simon Peter

文献摘要

相似文献

大容量非易失性存储器(NVM)是一种新的主存储器层。分层DRAM+NVM服务器可将总内存容量增加多达8倍,但如果管理不善,可能会将内存带宽减少多达7倍,并将延迟增加多达63%。我们研究了现有的硬件和软件分层内存管理系统上最近可用的英特尔Optane DC NVM与大数据应用程序,并发现没有现有的系统最大限度地提高应用程序的性能在真实的NVM。基于我们的研究结果,我们提出了HeMem,一个分层的主存管理系统从头开始设计的商业可用的NVM和使用它的大数据应用程序。HeMem管理分层的内存异步,存储和摊销内存访问跟踪,迁移,和相关的TLB同步开销。HeMem通过CPU事件(而不是页表)对内存访问进行采样,从而监控应用程序的内存使用情况。这使得HeMem能够扩展到TB级的内存,在快速内存中保持小而短暂的数据结构,并根据访问模式分配稀缺的非对称NVM带宽。最后,HeMem通过将每个应用程序的内存管理策略放置在用户级来实现灵活性。在采用英特尔Optane DC NVM的系统上,HeMem的性能优于基于硬件、操作系统和PL的分层内存管理,可将差距图形处理基准测试的运行时间减少多达50%,将筒仓内存数据库上的TPC-C吞吐量提高13%,将键值存储的性能隔离下的尾部延迟降低16%,并将NVM磨损减少多达10倍,而无需应用程序修改。
High-capacity non-volatile memory (NVM) is a new main memory tier. Tiered DRAM+NVM servers increase total memory capacity by up to 8x, but can diminish memory bandwidth by up to 7x and inflate latency by up to 63% if not managed well. We study existing hardware and software tiered memory management systems on the recently available Intel Optane DC NVM with big data applications and find that no existing system maximizes application performance on real NVM. Based on our findings, we present HeMem, a tiered main memory management system designed from scratch for commercially available NVM and the big data applications that use it. HeMem manages tiered memory asynchronously, batching and amortizing memory access tracking, migration, and associated TLB synchronization overheads. HeMem monitors application memory use by sampling memory access via CPU events, rather than page tables. This allows HeMem to scale to terabytes of memory, keeping small and ephemeral data structures in fast memory, and allocating scarce, asymmetric NVM bandwidth according to access patterns. Finally, HeMem is flexible by placing per-application memory management policy at user-level. On a system with Intel Optane DC NVM, HeMem outperforms hardware, OS, and PL-based tiered memory management, providing up to 50% runtime reduction for the GAP graph processing benchmark, 13% higher throughput for TPC-C on the Silo in-memory database, 16% lower tail-latency under performance isolation for a key-value store, and up to 10x less NVM wear than the next best solution, without application modification.