EMISSARY: Enhanced Miss Awareness Replacement Policy for L2 Instruction Caching

EMISSARY: Enhanced Miss Awareness Replacement Policy for L2 Instruction Caching
复制标题

DOI:
10.1145/3579371.3589097
复制
发表时间:
2023-06
期刊:
Proceedings of the 50th Annual International Symposium on Computer Architecture
影响因子:
--
通讯作者:
N. P. Nagendra;Bhargav Reddy Godala;Ishita Chaturvedi;Atmn Patel;Svilen Kanev;Tipp Moseley;Jared Stark;Gilles A. Pokam;Simone Campanoni;David I. August
N. P. Nagendra;Bhargav Reddy Godala;Ishita Chaturvedi;Atmn Patel;Svilen Kanev;Tipp Moseley;Jared Stark;Gilles A. Pokam;Simone Campanoni;David I. August
中科院分区:
其他
文献类型:
--
作者:
N. P. Nagendra;Bhargav Reddy Godala;Ishita Chaturvedi;Atmn Patel;Svilen Kanev;Tipp Moseley;Jared Stark;Gilles A. Pokam;Simone Campanoni;David I. August

文献摘要

被引文献

相似文献

几十年来,建筑师一直设计了缓存替换策略,以减少缓存失误。由于并非所有的缓存失误都同样影响处理器的性能,因此研究人员还提出了旨在降低总错过成本而不是总数遗漏的替换政策。但是,所有先前的成本感知替代政策均专门用于数据缓存,并且对于指导缓存是不合适或不必要的复杂。本文介绍了《使徒》,这是专门为教学缓存而设计的第一个成本感知的缓存替代家族。观察到现代体系结构完全容忍许多指令的缓存失误,使者抵制了驱逐出来的高速缓存线,这些高速公路的错过会导致昂贵的解码饥饿。在现代处理器的背景下,带有提取指导的预摘要和其他侵略性前端功能,适用于L2高速缓存指令的使者提供了令人印象深刻的3.24%的Geomean Speedup(最高23.7%),而Geomean Energy Savings则提供2.1%(上升到2.1%)当对具有大型代码足迹的广泛使用的服务器应用程序进行评估时,至17.7%。该加速度是无法实现的L2缓存获得的总速度的21.6%,其中所有能力和冲突指令都没有零周期的延迟延迟。
For decades, architects have designed cache replacement policies to reduce cache misses. Since not all cache misses affect processor performance equally, researchers have also proposed cache replacement policies focused on reducing the total miss cost rather than the total miss count. However, all prior cost-aware replacement policies have been proposed specifically for data caching and are either inappropriate or unnecessarily complex for instruction caching. This paper presents EMISSARY, the first cost-aware cache replacement family of policies specifically designed for instruction caching. Observing that modern architectures entirely tolerate many instruction cache misses, EMISSARY resists evicting those cache lines whose misses cause costly decode starvations. In the context of a modern processor with fetch-directed instruction prefetching and other aggressive front-end features, EMISSARY applied to L2 cache instructions delivers an impressive 3.24% geomean speedup (up to 23.7%) and a geomean energy savings of 2.1% (up to 17.7%) when evaluated on widely used server applications with large code footprints. This speedup is 21.6% of the total speedup obtained by an unrealizable L2 cache with a zero-cycle miss latency for all capacity and conflict instruction misses.