A Simple Pattern-matching Algorithm for Recovering Empty Nodes and their Antecedents

A Simple Pattern-matching Algorithm for Recovering Empty Nodes and their Antecedents
复制标题

DOI:
10.3115/1073083.1073107
复制
发表时间:
2002-07
期刊:
--
影响因子:
--
通讯作者:
Mark Johnson
Mark Johnson
中科院分区:
其他
文献类型:
--
作者:
Mark Johnson

文献摘要

被引文献

相似文献

本文描述了一种简单的模式匹配算法,用于恢复空节点,并在不包含此信息的短语结构树中识别它们的共同索引先行词。这些模式是最小的连接树片段,其中包含一个空节点以及与其共同索引的所有其他节点。本文还提出了一种与短语结构的大部分细节无关的空节点恢复过程的评估过程,从而可以将空节点恢复在句法分析器输出上的性能与黄金标准语料库中的空节点注释进行比较。在Charniak的解析器(Charniak,2000)和Penn Treebank(Marcus等人,1993)的输出上对该算法的评估表明,模式匹配算法在最频繁出现的空节点类型上表现得出人意料地好,因为它很简单。
This paper describes a simple pattern-matching algorithm for recovering empty nodes and identifying their co-indexed antecedents in phrase structure trees that do not contain this information. The patterns are minimal connected tree fragments containing an empty node and all other nodes co-indexed with it. This paper also proposes an evaluation procedure for empty node recovery procedures which is independent of most of the details of phrase structure, which makes it possible to compare the performance of empty node recovery on parser output with the empty node annotations in a gold-standard corpus. Evaluating the algorithm on the output of Charniak's parser (Charniak, 2000) and the Penn treebank (Marcus et al., 1993) shows that the pattern-matching algorithm does surprisingly well on the most frequently occuring types of empty nodes given its simplicity.