Stop Crying Over Your Cache Miss Rate: Handling Efficiently Thousands of Outstanding Misses in FPGAs

Stop Crying Over Your Cache Miss Rate: Handling Efficiently Thousands of Outstanding Misses in FPGAs
复制标题

别再为缓存缺失率而哭泣:高效处理 FPGA 中数以千计的未命中缺失

DOI:
--
复制
发表时间:
2019
期刊:
Symposium on Field Programmable Gate Arrays
影响因子:
--
通讯作者:
P. Ienne
P. Ienne
中科院分区:
--
文献类型:
--
作者:
Mikhail Asiatici;P. Ienne

文献摘要

被引文献

相似文献

即使时钟频率较低,FPGA 也依赖大规模数据路径并行性来加速应用程序。然而,稀疏线性代数和图形分析等应用程序的吞吐量受到对外部存储器的不规则访问的限制,对于外部存储器,由于非常频繁的未命中,典型的缓存几乎没有提供什么好处。非阻塞缓存在CPU上被广泛使用,以减少未命中的负面影响,从而提高缓存命中率较低的应用程序的性能;然而,它们依靠关联查找来处理多个未完成的未命中,这限制了它们的可扩展性,尤其是在 FPGA 上。当应用程序的命中率非常低时,这会导致频繁的停顿。在本文中,我们表明,通过在不停止的情况下处理数千个未命中的未命中,我们可以实现内存级并行性的大幅增加,这可以显着加速不规则的内存限制的延迟不敏感的应用程序。通过将未命中信息存储在 Block RAM(而不是关联内存)中的 Cuckoo 哈希表中,我们展示了如何修改非阻塞缓存以支持最多三个数量级的未命中。由此产生的未优化架构在面积延迟空间中为 12 个大型稀疏矩阵向量乘法基准提供了新的帕累托最优甚至帕累托主导设计点,与传统的命中优化架构相比,在面积减少 24 倍的情况下提供高达 25% 的加速,或者在类似面积的情况下提供高达 2 倍的加速。
FPGAs rely on massive datapath parallelism to accelerate applications even with a low clock frequency. However, applications such as sparse linear algebra and graph analytics have their throughput limited by irregular accesses to external memory for which typical caches provides little benefit because of very frequent misses. Non-blocking caches are widely used on CPUs to reduce the negative impact of misses and thus increase performance of applications with low cache hit rate; however, they rely on associative lookup for handling multiple outstanding misses, which limits their scalability, especially on FPGAs. This results in frequent stalls whenever the application has a very low hit rate. In this paper, we show that by handling thousands of outstanding misses without stalling we can achieve a massive increase of memory-level parallelism, which can significantly speed up irregular memory-bound latency-insensitive applications. By storing miss information in cuckoo hash tables in block RAM instead of associative memory, we show how a non-blocking cache can be modified to support up to three orders of magnitude more misses. The resulting miss-optimized architecture provides new Pareto-optimal and even Pareto-dominant design points in the area-delay space for twelve large sparse matrix-vector multiplication benchmarks, providing up to 25% speedup with 24x area reduction or to 2x speedup with similar area compared to traditional hit-optimized architectures.