SpZip: Architectural Support for Effective Data Compression In Irregular Applications

SpZip: Architectural Support for Effective Data Compression In Irregular Applications
复制标题

DOI:
10.1109/isca52012.2021.00087
复制
发表时间:
2021-06
期刊:
2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Yifan Yang;J. Emer;Daniel Sánchez
Yifan Yang;J. Emer;Daniel Sánchez
中科院分区:
其他
文献类型:
--
作者:
Yifan Yang;J. Emer;Daniel Sánchez

文献摘要

被引文献

相似文献

不规则的应用,如图分析和稀疏线性代数,表现出对单个或短序列元素的频繁的间接、依赖于数据的访问,这会导致高主存流量并限制性能。数据压缩是通过减少内存流量来加速非常规应用程序的一种很有前途的方法。然而,软件压缩增加了大量的开销,而现有的硬件压缩技术对非规则应用程序的复杂访问模式效果不佳,提出了一种使数据压缩适用于非规则算法的体系结构方法SpZip。SpZip加速了非常规应用程序使用的数据结构的遍历、解压缩和压缩。此外,这些活动以分离的方式运行,隐藏了内存访问和解压缩延迟。为了在这些应用中支持广泛的访问模式,SpZip是可编程的,并使用一种新颖的数据流配置语言来指定遍历和生成压缩数据的程序。我们的SpZip实现利用数据流执行和时分复用来实现廉价的可编程性。我们在一个模拟的多核系统上对SpZip进行了评估,该系统运行一系列广泛的图和线性代数算法。SpZip比以前最先进的纯软件(硬件加速)系统的性能平均提高了3.0×(1.5倍),内存流量减少了1.7×(1.4倍)。这些好处既来自减少因压缩而引起的数据移动,也源于减轻昂贵的遍历和(解)压缩操作的负担。
Irregular applications, such as graph analytics and sparse linear algebra, exhibit frequent indirect, data-dependent accesses to single or short sequences of elements that cause high main memory traffic and limit performance. Data compression is a promising way to accelerate irregular applications by reducing memory traffic. However, software compression adds substantial overheads, and prior hardware compression techniques work poorly on the complex access patterns of irregular applications.We present SpZip, an architectural approach that makes data compression practical for irregular algorithms. SpZip accelerates the traversal, decompression, and compression of the data structures used by irregular applications. In addition, these activities run in a decoupled fashion, hiding both memory access and decompression latencies. To support the wide range of access patterns in these applications, SpZip is programmable, and uses a novel Dataflow Configuration Language to specify programs that traverse and generate compressed data. Our SpZip implementation leverages dataflow execution and time-multiplexing to implement programmability cheaply. We evaluate SpZip on a simulated multicore system running a broad set of graph and linear algebra algorithms. SpZip outperforms prior state-of-the art software-only (hardware-accelerated) systems by gmean 3.0× (1.5×) and reduces memory traffic by 1.7× (1.4×). These benefits stem from both reducing data movement due to compression, and offloading expensive traversal and (de)compression operations.