Left-corner Parsing for Dependency Grammar

Left-corner Parsing for Dependency Grammar
复制标题

DOI:
10.5715/jnlp.22.251
复制
发表时间:
2015-12
影响因子:
--
通讯作者:
Hiroshi Noji;Yusuke Miyao
Hiroshi Noji;Yusuke Miyao
中科院分区:
--
文献类型:
--
作者:
Hiroshi Noji;Yusuke Miyao

文献摘要

相似文献

在这篇文章中,我们提出了一种增量依赖分析算法,该算法采用了左角分析策略的弧急切变体。我们的算法的堆栈深度捕捉到了识别的依赖结构的中心嵌入性。只有当处理较深的中心嵌入句子时,才会出现较高的堆栈深度,在这些句子中,人们发现难以理解。我们通过跨越19种语言的树库的两种实验来检验我们的算法是否能够捕捉到语言中普遍存在的句法规则。我们首先通过Oracle解析实验表明,与跨语言的其他算法相比,我们的解析算法识别带注释的树所需的堆栈深度始终较小。这一结果也表明了句法普遍性的存在,即深层中心嵌入是跨语言的罕见结构,这一结果尚未得到跨语言的定量检验。我们通过有监督的句法分析实验进一步研究了上述结论,结果表明,当跨语言解码时,我们提出的解析器对堆栈深度界限的约束始终不那么敏感,而其他解析器(如渴望弧形的解析器)的性能在很大程度上受到此类约束的影响。因此,我们得出结论,与现有解析器相比,我们解析器的堆栈深度代表了一种更有意义的衡量标准,可以用来捕获语言中的句法规则。
In this article, we present an incremental dependency parsing algorithm with an arc-eager variant of the left-corner parsing strategy. Our algorithm’s stack depth captures the center-embeddedness of the recognized dependency structure. A higher stack depth occurs only when processing deeper center-embedded sentences in which people find difficulty in comprehension. We examine whether our algorithm can capture the syntactic regularity that universally exists in languages through two kinds of experiments across treebanks of 19 languages. We first show through oracle parsing experiments that our parsing algorithm consistently requires less stack depth to recognize annotated trees relative to other algorithms across languages. This result also suggests the existence of a syntactic universal by which deeper center-embedding is a rare construction across languages, a result that has yet to be quantitatively cross-linguistically examined. We further investigate the above claim through supervised parsing experiments and show that our proposed parser is consistently less sensitive to constraints on stack depth bounds when decoding across languages, while the performance of other parsers such as the arc-eager parser is largely affected by such constraints. We thus conclude that the stack depth of our parser represents a more meaningful measure for capturing syntactic regularity in languages than those of existing parsers.