Using Linguistic Principles to Recover Empty Categories

Using Linguistic Principles to Recover Empty Categories
复制标题

DOI:
10.3115/1218955.1219037
复制
发表时间:
2004-07
期刊:
--
影响因子:
--
通讯作者:
R. Campbell
R. Campbell
中科院分区:
其他
文献类型:
--
作者:
R. Campbell

文献摘要

被引文献

相似文献

本文描述了一种算法,用于检测 Penn Treebank 中的空节点(Marcus 等人,1993),找到它们的前因,并为它们分配功能标签,而无需访问诸如效价之类的词汇信息。与之前完成此任务的方法不同,当前的方法不是基于语料库的,而是利用了早期政府约束理论(Chomsky,1981)的原则,即作为注释基础的句法理论。使用 Johnson (2002) 提出的评估指标,在给定带注释的输入剥离空类别或解析器的输出的情况下,该方法在空类别检测和先行词识别方面均优于先前发布的方法。指出了该评估指标的一些问题,并根据结果提出了替代方案。本文考虑了解决该问题的基于原则的方法应优于基于语料库的方法的原因,并推测了混合方法的可能性。
This paper describes an algorithm for detecting empty nodes in the Penn Treebank (Marcus et al., 1993), finding their antecedents, and assigning them function tags, without access to lexical information such as valency. Unlike previous approaches to this task, the current method is not corpus-based, but rather makes use of the principles of early Government-Binding theory (Chomsky, 1981), the syntactic theory that underlies the annotation. Using the evaluation metric proposed by Johnson (2002), this approach outperforms previously published approaches on both detection of empty categories and antecedent identification, given either annotated input stripped of empty categories or the output of a parser. Some problems with this evaluation metric are noted and an alternative is proposed along with the results. The paper considers the reasons a principle-based approach to this problem should outperform corpus-based approaches, and speculates on the possibility of a hybrid approach.