Cache automaton

Cache automaton
复制标题

缓存自动机

DOI:
10.1145/3123939.3123986
复制
发表时间:
2017
期刊:
MICRO-50 '17 Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Das, Reetuparna
Das, Reetuparna
中科院分区:
--
文献类型:
--
作者:
Subramaniyan, Arun;Wang, Jingcheng;Balasubramanian, Ezhil R.;Blaauw, David;Sylvester, Dennis;Das, Reetuparna

文献摘要

相似文献

有限状态自动机在DNA测序和XML解析等新兴应用领域中被广泛用于加速模式匹配。传统的CPU和以计算为中心的加速器被自动机处理中的内存带宽和不规则的内存访问模式所困扰,我们提出了Cache Automaton,它重新利用自动机处理的末级缓存,以及一个编译器,它自动化将大型真实的世界非确定性有限自动机(NFA)映射到所提出的架构的过程。缓存自动机扩展了传统的末级缓存架构,组件加速NFA处理的两个阶段:状态匹配和状态转换。使用感测放大器循环技术,利用符号匹配的空间局部性,使状态匹配有效。使用一种新的紧凑型交换机架构的状态转换效率。通过重叠这两个阶段的相邻符号,我们实现了一个有效的流水线设计。我们评估两个设计,一个优化的性能和其他优化的空间,在一组20个不同的基准。性能优化的设计提供了一个15倍的加速比基于DRAM的美光的自动机处理器和3840倍的加速比在传统的x86 CPU的处理。所提出的设计在基准测试中平均利用1.2MB的缓存空间,而每个输入符号消耗2.3nJ的能量。我们的空间优化设计可以将该高速缓存利用率降低到0.72 MB,同时仍然提供比AP高9倍的加速比。
Finite State Automata are widely used to accelerate pattern matching in many emerging application domains like DNA sequencing and XML parsing. Conventional CPUs and compute-centric accelerators are bottlenecked by memory bandwidth and irregular memory access patterns in automata processing.We presentCache Automaton, which repurposes last-level cache for automata processing, and a compiler that automates the process of mapping large real world Non-Deterministic Finite Automata (NFAs) to the proposed architecture. Cache Automaton extends a conventional last-level cache architecture with components to accelerate two phases in NFA processing: state-match and state-transition. State-matching is made efficient using a sense-amplifier cycling technique that exploits spatial locality in symbol matches. State-transition is made efficient using a new compact switch architecture. By overlapping these two phases for adjacent symbols we realize an efficient pipelined design.We evaluate two designs, one optimized for performance and the other optimized for space, across a set of 20 diverse benchmarks. The performance optimized design provides a speedup of 15× over DRAM-based Micron's Automata Processor and 3840× speedup over processing in a conventional x86 CPU. The proposed design utilizes on an average 1.2MBof cache space across benchmarks, while consuming 2.3nJof energy per input symbol. Our space optimized design can reduce the cache utilization to 0.72MB, while still providing a speedup of 9× over AP.