A Formal Analysis of the NVIDIA PTX Memory Consistency Model

A Formal Analysis of the NVIDIA PTX Memory Consistency Model
复制标题

NVIDIA PTX 内存一致性模型的形式分析

DOI:
10.1145/3297858.3304043
复制
发表时间:
2019
期刊:
Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Olivier Giroux
Olivier Giroux
中科院分区:
--
文献类型:
--
作者:
Daniel Lustig;Sameer D. Sahasrabuddhe;Olivier Giroux

文献摘要

参考文献

被引文献

相似文献

本文首次对NVIDIA PTX虚拟ISA的官方内存一致性模型进行了形式化分析。像其他GPU内存模型一样,PTX内存模型是弱有序的,但提供了作用域同步原语,使GPU程序线程能够通过内存进行通信。然而,与一些竞争的GPU内存模型不同,PTX不需要数据竞争自由,这导致PTX在其内存模型中使用了一套完全不同(也更复杂)的规则。因此,PTX显然需要一个严格可靠的内存模型测试和分析基础设施。我们将PTX内存模型的形式化分析分解为多个步骤,共同证明其严谨性和有效性。首先,我们将公共PTX文档中的英语语言规范改编为正式的公理模型。其次,我们导出了一个类似于opencl的作用域c++模型的最新表示,并开发了一个从该作用域c++模型的同步原语到PTX的映射。第三,使用Alloy关系建模工具对映射的正确性进行实证检验。最后,我们将模型和映射编译到Coq中,并构建了一个完整的机器检查证明,证明映射对于任何大小的程序都是可靠的。我们的分析表明,尽管前几代存在问题,但新的NVIDIA PTX内存模型适合作为GPU编程语言(如CUDA)的声音编译目标。
This paper presents the first formal analysis of the official memory consistency model for the NVIDIA PTX virtual ISA. Like other GPU memory models, the PTX memory model is weakly ordered but provides scoped synchronization primitives that enable GPU program threads to communicate through memory. However, unlike some competing GPU memory models, PTX does not require data race freedom, and this results in PTX using a fundamentally different (and more complicated) set of rules in its memory model. As such, PTX has a clear need for a rigorous and reliable memory model testing and analysis infrastructure. We break our formal analysis of the PTX memory model into multiple steps that collectively demonstrate its rigor and validity. First, we adapt the English language specification from the public PTX documentation into a formal axiomatic model. Second, we derive an up-to-date presentation of an OpenCL-like scoped C++ model and develop a mapping from the synchronization primitives of that scoped C++ model onto PTX. Third, we use the Alloy relational modeling tool to empirically test the correctness of the mapping. Finally, we compile the model and mapping into Coq and build a full machine-checked proof that the mapping is sound for programs of any size. Our analysis demonstrates that in spite of issues in previous generations, the new NVIDIA PTX memory model is suitable as a sound compilation target for GPU programming languages such as CUDA.
彻底修改 C11 和 OpenCL 中的 SC 原子
DOI: 10.1145/2837614.2837637
发表时间: 2016
期刊: --
影响因子: --
作者:
Batty M
通讯作者: Batty M
自动比较内存一致性模型
DOI: 10.1145/3009837.3009838
发表时间: 2017
期刊: --
影响因子: --
作者:
Wickerson J
通讯作者: Wickerson J
DOI: 10.1145/3276506
发表时间: 2018
影响因子: --
作者:
Ou, Peizhao;Demsky, Brian
通讯作者: Demsky, Brian
混合大小并发:ARM、POWER、C/C 11 和 SC
DOI: 10.1145/3009837.3009839
发表时间: 2017
期刊: --
影响因子: --
作者:
Flur S
通讯作者: Flur S
了解 POWER 多处理器
DOI: 10.1145/1993316.1993520
发表时间: 2011
影响因子: --
作者:
Sarkar S
通讯作者: Sarkar S