SemCache: semantics-aware caching for efficient GPU offloading

SemCache: semantics-aware caching for efficient GPU offloading
复制标题

SemCache:语义感知缓存,可实现高效 GPU 卸载

DOI:
10.1145/2464996.2465021
复制
发表时间:
2016
期刊:
J. Adv. Comput. Intell. Intell. Informatics
影响因子:
--
通讯作者:
Milind Kulkarni
Milind Kulkarni
中科院分区:
--
文献类型:
--
作者:
Nabeel AlSaber;Milind Kulkarni

文献摘要

被引文献

相似文献

最近,GPU 库通过将计算卸载到 GPU 来轻松提高应用程序性能。然而,使用此类库会带来手动处理 GPU 和 CPU 内存空间之间显式数据移动的复杂性。不幸的是,当在复杂的应用程序中使用这些库时,很难优化多个内核调用之间的 CPU-GPU 通信以避免冗余通信。 在本文中,我们介绍了 SemCache,这是一种语义感知的 GPU 缓存,它可以自动管理 CPU-GPU 通信,并通过使用缓存消除冗余传输来动态优化通信。其主要功能是使用库语义来确定给定卸载库(例如 BLAS 中的矩阵)的适当缓存粒度。我们将 SemCache 应用于 BLAS 库,以提供一个 GPU 直接替换库,自动处理通信和优化。我们的缓存技术是高效的;它只跟踪矩阵,而不是细粒度地跟踪每个内存访问。实验结果表明,我们的系统可以显着减少现实计算科学应用程序的冗余通信,并提供显着的性能改进,击败基于 GPU 的实现(如 CULA 和 CUBLAS)。
Recently, GPU libraries have made it easy to improve application performance by offloading computation to the GPU. However, using such libraries introduces the complexity of manually handling explicit data movements between GPU and CPU memory spaces. Unfortunately, when using these libraries with complex applications, it is very difficult to optimize CPU-GPU communication between multiple kernel invocations to avoid redundant communication. In this paper, we introduce SemCache, a semantics-aware GPU cache that automatically manages CPU-GPU communication and dynamically optimizes communication by eliminating redundant transfers using caching. Its key feature is the use of library semantics to determine the appropriate caching granularity for a given offloaded library (e.g., matrices in BLAS). We applied SemCache to BLAS libraries to provide a GPU drop-in replacement library which handles communications and optimizations automatically. Our caching technique is efficient; it only tracks matrices instead of tracking every memory access at fine granularity. Experimental results show that our system can dramatically reduce redundant communication for real-world computational science application and deliver significant performance improvements, beating GPU-based implementations like CULA and CUBLAS.