A Simple Pattern-matching Algorithm for Recovering Empty Nodes and their Antecedents
A Simple Pattern-matching Algorithm for Recovering Empty Nodes and their Antecedents
复制标题
DOI:
10.3115/1073083.1073107
复制
发表时间:
2002-07
期刊:
影响因子:
--
通讯作者:
Mark Johnson
中科院分区:
文献类型:
--
作者:
Mark Johnson
This paper describes a simple pattern-matching algorithm for recovering empty nodes and identifying their co-indexed antecedents in phrase structure trees that do not contain this information. The patterns are minimal connected tree fragments containing an empty node and all other nodes co-indexed with it. This paper also proposes an evaluation procedure for empty node recovery procedures which is independent of most of the details of phrase structure, which makes it possible to compare the performance of empty node recovery on parser output with the empty node annotations in a gold-standard corpus. Evaluating the algorithm on the output of Charniak's parser (Charniak, 2000) and the Penn treebank (Marcus et al., 1993) shows that the pattern-matching algorithm does surprisingly well on the most frequently occuring types of empty nodes given its simplicity.