Romulus: Efficient Algorithms for Persistent Transactional Memory

Romulus: Efficient Algorithms for Persistent Transactional Memory
复制标题

Romulus:持久事务内存的高效算法

DOI:
10.1145/3210377.3210392
复制
发表时间:
2018
期刊:
Proceedings of the 30th on Symposium on Parallelism in Algorithms and Architectures
影响因子:
--
通讯作者:
P. Ramalhete
P. Ramalhete
中科院分区:
--
文献类型:
--
作者:
Andreia Correia;P. Felber;P. Ramalhete

文献摘要

被引文献

相似文献

字节可寻址持久性存储器消除了对数据的串行化和非串行化的需要,并从持久性存储器,允许应用程序通过公共存储和加载指令与它交互。在进程或系统故障的情况下,应用程序依赖于持久性技术来在非易失性存储器(NVM)中提供一致的数据存储。对于这些技术中的大多数,通过更新的日志记录来确保一致性,随之而来的是密集的高速缓存行刷新和持久性围栏,以保证正确性。基于撤销日志的方法需要在每次就地修改之前存储插入和持久化栅栏。基于重做日志的技术可以只使用两个持久性围栏来执行事务,尽管它们需要存储和加载插入,这可能会导致大型事务的性能损失。到目前为止,这些技术难以与已知的存储器分配器集成,需要专门为NVM设计的分配器或垃圾收集器。我们提出了Romulus,一个用户级库持久性事务存储器(PTM),它通过使用数据的双副本提供持久的事务。Romulus中的事务最多需要四个持久性围栏,而不管事务大小如何。罗穆卢斯只使用存储插入。内存分配器的任何顺序实现都可以适用于Romulus。由于其轻量级设计和低同步开销,Romulus在仅更新工作负载中实现了当前最先进PTM的两倍吞吐量,并且在主要读取场景中实现了一个数量级以上。
Byte addressable persistent memory eliminates the need for serialization and deserialization of data, to and from persistent storage, allowing applications to interact with it through common store and load instructions. In the event of a process or system failure, applications rely on persistent techniques to provide consistent storage of data in non-volatile memory (NVM). For most of these techniques, consistency is ensured through logging of updates, with consequent intensive cache line flushing and persistent fences necessary to guarantee correctness. Undo log based approaches require store interposition and persistence fences before each in-place modification. Redo log based techniques can execute transactions using just two persistence fences, although they require store and load interposition which may incur a performance penalty for large transactions. So far, these techniques have been difficult to integrate with known memory allocators, requiring allocators or garbage collectors specifically designed for NVM. We present Romulus, a user-level library persistent transactional memory (PTM) which provides durable transactions through the usage of twin copies of the data. A transaction in Romulus requires at most four persistence fences, regardless of the transaction size. Romulus uses only store interposition. Any sequential implementation of a memory allocator can be adapted to work with Romulus. Thanks to its lightweight design and low synchronization overhead, Romulus achieves twice the throughput of current state of the art PTMs in update-only workloads, and more than one order of magnitude in read-mostly scenarios.