Brill tagging on the Micron Automata Processor

Brill tagging on the Micron Automata Processor
复制标题

Micron Automata 处理器上的 Brill 标记

DOI:
--
复制
发表时间:
2015
期刊:
Proceedings of the 2015 IEEE 9th International Conference on Semantic Computing (IEEE ICSC 2015)
影响因子:
--
通讯作者:
K. Skadron
K. Skadron
中科院分区:
--
文献类型:
--
作者:
Keira Zhou;J. J. Fox;Ke Wang;Donald E. Brown;K. Skadron

文献摘要

被引文献

相似文献

语义分析通常使用一系列自然语言处理 (NLP) 工具,例如词性 (POS) 标记。 Brill 标记是 NLP 中基于规则的经典 POS 标记算法。然而,在传统的冯诺依曼架构上,标记器的实现本质上很慢。在本文中,我们在 Micron Automata 处理器上加速了 Brill 标记的第二阶段,这是一种可以并行执行大规模模式匹配的新计算架构。所设计的结构使用 218 条上下文规则通过布朗语料库的子集进行测试。结果显示,与 CPU 上的单线程实现相比,在单个 AP 芯片上实现的第二级标记器的速度提高了 38 倍。这种加速与规则的数量成线性关系,从而使得大型和/或复杂的规则集在计算上变得实用。本文介绍了这种新的加速器在计算语言任务中的用途,特别是那些涉及基于规则或模式匹配方法的任务。
Semantic analysis often uses a pipeline of Natural Language Processing (NLP) tools such as part-of-speech (POS) tagging. Brill tagging is a classic rule-based algorithm for POS tagging within NLP. However, implementation of the tagger is inherently slow on conventional Von Neumann architectures. In this paper, we accelerate the second stage of Brill tagging on the Micron Automata Processor, a new computing architecture that can perform massive pattern matching in parallel. The designed structure is tested with a subset of the Brown Corpus using 218 contextual rules. The results show a 38X speed-up for the second stage tagger implemented on a single AP chip, compared to a single thread implementation on CPU. This speed-up is linear with the number of rules, thus making large and/or complex rule sets computationally practical. This paper introduces the use of this new accelerator for computational linguistic tasks, particularly those that involve rule-based or pattern-matching approaches.