A Simple Cache Coherence Scheme for Integrated CPU-GPU Systems

A Simple Cache Coherence Scheme for Integrated CPU-GPU Systems
复制标题

DOI:
10.1109/dac18072.2020.9218664
复制
发表时间:
2020-07
期刊:
2020 57th ACM/IEEE Design Automation Conference (DAC)
影响因子:
--
通讯作者:
Ardhi Wiratama Baskara Yudha;Reza Pulungan;H. Hoffmann;Yan Solihin
Ardhi Wiratama Baskara Yudha;Reza Pulungan;H. Hoffmann;Yan Solihin
中科院分区:
其他
文献类型:
--
作者:
Ardhi Wiratama Baskara Yudha;Reza Pulungan;H. Hoffmann;Yan Solihin

文献摘要

相似文献

本文提出了一种新的方法来加速在集成的CPU-GPU系统上运行的应用程序。许多集成的CPU-GPU系统使用高速缓存一致性共享内存进行通信。例如,在CPU产生用于GPU的数据之后,GPU可以在其访问数据时将数据拉入其高速缓存中。在这种基于拉取的方法中,数据驻留在共享高速缓存中,直到GPU访问它,导致在第一次GPU访问高速缓存行时的长加载延迟。在这项工作中,我们提出了一种新的,基于推的,一致性机制,明确利用CPU和GPU的生产者-消费者的关系,自动移动数据从CPU到GPU的最后一级缓存。所提出的机制导致GPU L2高速缓存未命中率总体上显著降低,从而提高整体性能。我们的实验表明,该方案可以提高性能高达37%,与典型的改善5-7%的范围。我们发现,即使当测试的应用程序不受益于所提出的方法,他们的性能不会降低与我们的技术。虽然我们演示了如何提出的计划可以与传统的缓存一致性机制共存,我们认为,它也可以被用来作为一个更简单的替代现有的协议。
This paper presents a novel approach to accelerate applications running on integrated CPU-GPU systems. Many integrated CPU-GPU systems use cache-coherent shared memory to communicate. For example, after CPU produces data for GPU, the GPU may pull the data into its cache when it accesses the data. In such a pull-based approach, data resides in a shared cache until the GPU accesses it, resulting in long load latency on a first GPU access to a cache line. In this work, we propose a new, push-based, coherence mechanism that explicitly exploits the CPU and GPU producer-consumer relationship by automatically moving data from CPU to GPU last-level cache. The proposed mechanism results in a dramatic reduction of the GPU L2 cache miss rate in general, and a consequent increase in overall performance. Our experiments show that the proposed scheme can increase performance by up to 37%, with typical improvements in the 5–7% range. We find that even when tested applications do not benefit from the proposed approach, their performance does not decrease with our technique. While we demonstrate how the proposed scheme can co-exist with traditional cache coherence mechanisms, we argue that it could also be used as a simpler replacement for existing protocols.