Perceptron-Based Prefetch Filtering

Perceptron-Based Prefetch Filtering
复制标题

DOI:
10.1145/3307650.3322207
复制
发表时间:
2019-06
期刊:
2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Eshan Bhatia;Gino Chacon;Seth H. Pugsley;Elvira Teran;Paul V. Gratz;Daniel A. Jiménez
Eshan Bhatia;Gino Chacon;Seth H. Pugsley;Elvira Teran;Paul V. Gratz;Daniel A. Jiménez
中科院分区:
其他
文献类型:
--
作者:
Eshan Bhatia;Gino Chacon;Seth H. Pugsley;Elvira Teran;Paul V. Gratz;Daniel A. Jiménez

文献摘要

相似文献

在现代处理器设计中,硬件预取是隐藏高速缓存未命中延迟的有效技术。预取器性能可以由两个通常彼此不一致的主要度量来表征:覆盖率,预取器带入该高速缓存的基线缓存未命中的分数;以及准确度,最终使用的预取的分数。过度激进的预取器可能会以降低准确性为代价来提高覆盖率。因此,性能可能会受到这种过度攻击的损害,因为浪费了许多资源,包括高速缓存容量和带宽。一个理想的预取器应该同时具有高覆盖率和准确性。在本文中,我们介绍了基于感知器的预取过滤(PPF)作为一种方法,以增加覆盖率的预取生成的底层预取,而不会产生负面影响的准确性。PPF支持对底层预取器进行更积极的调优,通过过滤掉越来越多的不准确预取来增加覆盖率。我们还探索了一系列用于训练PPF感知器层的功能,以识别不准确的预取。PPF在SPEC CPU 2017基准测试的内存密集型子集上的性能提高了3.78%(单核配置)和11.4%(4核配置),与单独的底层预取器相比。
Hardware prefetching is an effective technique for hiding cache miss latencies in modern processor designs. Prefetcher performance can be characterized by two main metrics that are generally at odds with one another: coverage, the fraction of baseline cache misses which the prefetcher brings into the cache; and accuracy, the fraction of prefetches which are ultimately used. An overly aggressive prefetcher may improve coverage at the cost of reduced accuracy. Thus, performance may be harmed by this over-aggressiveness because many resources are wasted, including cache capacity and bandwidth. An ideal prefetcher would have both high coverage and accuracy. In this paper, we introduce Perceptron-based Prefetch Filtering (PPF) as a way to increase the coverage of the prefetches generated by an underlying prefetcher without negatively impacting accuracy. PPF enables more aggressive tuning of the underlying prefetcher, leading to increased coverage by filtering out the growing numbers of inaccurate prefetches such an aggressive tuning implies. We also explore a range of features to use to train PPF's perceptron layer to identify inaccurate prefetches. PPF improves performance on a memory-intensive subset of the SPEC CPU 2017 benchmarks by 3.78% for a single-core configuration, and by 11.4% for a 4-core configuration, compared to the underlying prefetcher alone.