Remote Memory References at Block Granularity

Remote Memory References at Block Granularity
复制标题

块粒度的远程内存引用

DOI:
--
复制
发表时间:
2017
期刊:
International Conference on Principles of Distributed Systems
影响因子:
--
通讯作者:
Gili Yavneh
Gili Yavneh
中科院分区:
--
文献类型:
--
作者:
H. Attiya;Gili Yavneh

文献摘要

被引文献

相似文献

访问远程内存中存储的共享对象的成本,而忽略了对本地内存中的共享对象的访问,可以通过执行中的两个口味 - 远程内存引用(RMR)的数量来评估。 CACHE-COHERENT(CC)和分布式共享内存(DSM) - 模型两个流行的共享 - 内存架构。访问,即访问共享内存的事实是在块中执行的。 本文提出了一种称为块RMR的新度量,计算远程内存引用的数量,同时考虑到共享对象可以一方面分组为块。对本地内存的共享对象可能会保存另一个将另一个对象放置在同一块上的RMR。由于同时访问同一块中的另一个对象,访问对象。 本文证明,在CC和DSM模型中,当对象具有不同的尺寸,即使在CC模型中,找到最佳的位置是NP-HARD。当一个块可以存储三个或更多的对象时,即使在DSM模型中已知访问序列时,该障碍物也可以存储三个或更多。通知过程的机制不再有效,即,如果支持廉价的无效,则支持缓存相干。可以通过将每个对象放置在最常访问它的过程的内存中,如果访问的序列是事先知道的。
The cost of accessing shared objects that are stored in remote memory, while neglecting accesses to shared objects that are cached in the local memory, can be evaluated by the number of remote memory references (RMRs) in an execution. Two flavours of this measure—cache-coherent (CC) and distributed shared memory (DSM)—model two popular shared-memory architectures. The number of RMRs, however, does not take into account the granularity of memory accesses, namely, the fact that accesses to the shared memory are performed in blocks. This paper proposes a new measure, called block RMRs, counting the number of remote memory references while taking into account the fact that shared objects can be grouped into blocks. On the one hand, this measure reflects the fact that the RMR incurred for bringing a shared object to the local memory might save another RMR for bringing another object placed at the same block. On the other hand, this measure accounts for false sharing: the fact that an RMR may be incurred when accessing an object due to a concurrent access to another object in the same block. This paper proves that in both the CC and the DSM models, finding an optimal placement is NP-hard when objects have different sizes, even for two processes. In the CC model, finding an optimal placement, i.e., grouping of objects into blocks, is NP-hard when a block can store three objects or more; the result holds even if the sequence of accesses is known in advance. In the DSM model, the answer depends on whether there is an efficient mechanism to inform processes that the data in their local memory is no longer valid, i.e., cache coherence is supported. If coherence is supported with cheap invalidation, then finding an optimal solution is NP-hard. If coherence is not supported, an optimal placement can be achieved by placing each object in the memory of the process that accesses it most often, if the sequence of accesses is known in advance.