Lazy release consistency for GPUs

Lazy release consistency for GPUs
复制标题

GPU 的延迟释放一致性

DOI:
10.1109/micro.2016.7783729
复制
发表时间:
2016
期刊:
2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
D. Wood
D. Wood
中科院分区:
--
文献类型:
--
作者:
Johnathan Alsop;Marc S. Orr;Bradford M. Beckmann;D. Wood

文献摘要

参考文献

被引文献

相似文献

异构无竞争(HRF)内存模型已被异构系统架构(HSA)基金会和OpenCLTM所接受,因为它清晰而精确地定义了当前GPU的行为。然而,与更简单的SC for DRF存储器模型相比,HRF具有两个缺点。首先,HRF要求程序员用正确的同步范围来标记原子内存操作。这种显式标记可以节省显着的一致性开销时,同步是本地的,但它是繁琐和容易出错。第二个缺点是HRF限制了重要的动态数据共享模式,如工作窃取。远程作用域提升(RSP)的先前工作试图解决第二个缺点。然而,RSP进一步复杂化的内存模型,并没有可扩展的实现RSP已经提出。例如,我们发现,与放弃工作窃取和范围的天真基线系统相比,之前提出的RSP实现实际上会导致大型GPU上的速度下降高达30%。与此同时,DeNovo已被证明可以提供与SC for DRF内存模型的高效同步,平均性能比我们的基线系统高出21%,但它引入了额外的开销来维护所有修改数据的所有权。为了解决这些不足之处,我们建议适应懒惰释放一致性-以前只提出了同构CPU系统-异构系统。我们的方法,称为hLRC,使用DeNovo-like机制来跟踪同步变量的所有权,仅当同步变量改变位置时才懒洋洋地执行一致性操作。hLRC允许GPU程序员使用更简单的SC for DRF内存模型,而无需跟踪所有修改数据的所有权。我们的评估表明,延迟发布一致性在一组工作窃取图分析应用程序中提供了强大的性能改进-与基线系统相比平均提高29%。
The heterogeneous-race-free (HRF) memory model has been embraced by the Heterogeneous System Architecture (HSA) Foundation and OpenCLTM because it clearly and precisely defines the behavior of current GPUs. However, compared to the simpler SC for DRF memory model, HRF has two shortcomings. The first is that HRF requires programmers to label atomic memory operations with the correct scope of synchronization. This explicit labeling can save significant coherence overhead when synchronization is local, but it is tedious and error-prone. The second shortcoming is that HRF restricts important dynamic data sharing patterns like work stealing. Prior work on remote-scope promotion (RSP) attempted to resolve the second shortcoming. However, RSP further complicates the memory model and no scalable implementation of RSP has been proposed. For example, we found that the previously proposed RSP implementation actually results in slowdowns of up to 30% on large GPUs, compared to a naïve baseline system that forgoes work stealing and scopes. Meanwhile, DeNovo has been shown to offer efficient synchronization with an SC for DRF memory model, performing on average 21% better than our baseline system, but it introduces additional overheads to maintain ownership of all modified data. To resolve these deficiencies, we propose to adapt lazy release consistency - previously only proposed for homogeneous CPU systems - to a heterogeneous system. Our approach, called hLRC, uses a DeNovo-like mechanism to track ownership of synchronization variables, lazily performing coherence actions only when a synchronization variable changes locations. hLRC allows GPU programmers to use the simpler SC for DRF memory model without tracking ownership for all modified data. Our evaluation shows that lazy release consistency provides robust performance improvement across a set of work-stealing graph analysis applications - 29% on average versus the baseline system.
远程推广:澄清、整改、核实
DOI: 10.1145/2814270.2814283
发表时间: 2015
期刊: --
影响因子: --
作者:
Wickerson J
通讯作者: Wickerson J
RC3:具有 RC 扩展的 x86-64 一致性定向缓存一致性
DOI: 10.1109/pact.2015.37
发表时间: 2015
期刊: --
影响因子: --
作者:
Elver M
通讯作者: Elver M