Systems-on-Chip with Strong Ordering

Systems-on-Chip with Strong Ordering
复制标题

具有强排序功能的片上系统

DOI:
10.1145/3428153
复制
发表时间:
2021
影响因子:
1.6
通讯作者:
Lipasti, Mikko H.
Lipasti, Mikko H.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Puthoor, Sooraj;Lipasti, Mikko H.

文献摘要

参考文献

被引文献

相似文献

顺序一致性(SC)是最直观的内存一致性模型,也是程序员和硬件设计人员最容易推理的模型。然而,从性能的角度来看,SC所施加的严格的内存排序限制使其不那么有吸引力。此外,以前的高性能SC实现需要复杂的硬件结构来支持推测和recovery.In这篇文章中,我们介绍了锁步SC一致性模型(LSC),一个新的内存模型的基础上SC,但仔细定义,以适应数据并行锁步执行的GPU范例。我们还描述了一个有效的LSC实现的APU片上系统(SoC),并表明我们的实现执行接近基线放松模型。我们的实现的评估表明,几何平均性能成本锁步SC仅为0.76%的GPU执行和6.11%的整个APU SoC相比,基线与较弱的内存一致性模型。在未来的APU和SoC设计中采用LSC将减轻程序员编写正确并行程序的负担,同时还可以简化具有异构处理元件和复杂内存层次结构的系统的实现和验证。
Sequential consistency (SC) is the most intuitive memory consistency model and the easiest for programmers and hardware designers to reason about. However, the strict memory ordering restrictions imposed by SC make it less attractive from a performance standpoint. Additionally, prior high-performance SC implementations required complex hardware structures to support speculation and recovery.In this article, we introduce the lockstep SC consistency model (LSC), a new memory model based on SC but carefully defined to accommodate the data parallel lockstep execution paradigm of GPUs. We also describe an efficient LSC implementation for an APU system-on-chip (SoC) and show that our implementation performs close to the baseline relaxed model. Evaluation of our implementation shows that the geometric mean performance cost for lockstep SC is just 0.76% for GPU execution and 6.11% for the entire APU SoC compared to a baseline with a weaker memory consistency model. Adoption of LSC in future APU and SoC designs will reduce the burden on programmers trying to write correct parallel programs, while also simplifying the implementation and verification of systems with heterogeneous processing elements and complex memory hierarchies.1
氨纶:实现高效异质相干的灵活接口
DOI: --
发表时间: 2018
期刊: International Symposium on Computer Architecture
影响因子: --
作者:
Johnathan Alsop;Matthew D. Sinclair;S. Adve
通讯作者: S. Adve
使用推测性退休和更大的指令窗口来缩小内存一致性模型之间的性能差距
DOI: 10.1145/258492.258512
发表时间: 1997
期刊: Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques
影响因子: --
作者:
Parthasarathy Ranganathan;Vijay S. Pai;S. Adve
通讯作者: S. Adve
推测顺序一致性,几乎没有自定义存储
DOI: 10.1109/pact.2002.1106016
发表时间: 2002
期刊: Proceedings.International Conference on Parallel Architectures and Compilation Techniques
影响因子: --
作者:
C. Gniady;B. Falsafi
通讯作者: B. Falsafi
DOI: 10.1007/978-3-540-70545-1_12
发表时间: 2008-07
期刊: --
影响因子: --
作者:
S. Burckhardt;M. Musuvathi
通讯作者: S. Burckhardt;M. Musuvathi
DOI: 10.1145/2837614.2837615
发表时间: 2016-01
期刊: Proceedings of the 43rd Annual ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages
影响因子: --
作者:
Shaked Flur;Kathryn E. Gray;Christopher Pulte;Susmit Sarkar;A. Sezgin;Luc Maranget;Will Deacon;Peter Sewell
通讯作者: Shaked Flur;Kathryn E. Gray;Christopher Pulte;Susmit Sarkar;A. Sezgin;Luc Maranget;Will Deacon;Peter Sewell