Data-Oriented Parsing with Discontinuous Constituents and Function Tags

Data-Oriented Parsing with Discontinuous Constituents and Function Tags
复制标题

DOI:
10.15398/jlm.v4i1.100
复制
发表时间:
2016-04
期刊:
J. Lang. Model.
影响因子:
--
通讯作者:
Andreas van Cranenburgh;R. Scha;R. Bod
Andreas van Cranenburgh;R. Scha;R. Bod
中科院分区:
其他
文献类型:
--
作者:
Andreas van Cranenburgh;R. Scha;R. Bod

文献摘要

被引文献

相似文献

统计分析器是有效的,但通常限于产生投影依赖或成分。另一方面,语言丰富的解析器识别非局部关系并分析形式和功能现象,但依赖于广泛的手动语法开发。我们联合收割机两者的优势,建立一个统计分析器,产生更丰富的分析。我们研究新的技术来实现基于树库的解析器,允许不连续的成分。我们提出了两个系统。一个系统是基于字符串重写线性上下文无关重写系统(LCFRS),同时使用概率不连续树替换语法(PDTSG),以提高消歧性能。另一种系统对短语结构树的标签中的不连续性进行编码,从而允许高效的上下文无关语法解析。这两个系统表明,在树替换语法中使用的树片段提高了消歧性能,同时根据需要捕获非本地关系。此外,我们提出的模型,产生功能标签的结果,从而在语言上更充分的数据模型。我们报告了大量的准确性改进不连续解析德语,英语和荷兰语,包括荷兰语口语的结果。
Statistical parsers are e ective but are typically limited to producing projective dependencies or constituents. On the other hand, linguisti- cally rich parsers recognize non-local relations and analyze both form and function phenomena but rely on extensive manual grammar development. We combine advantages of the two by building a statistical parser that produces richer analyses. We investigate new techniques to implement treebank-based parsers that allow for discontinuous constituents. We present two systems. One system is based on a string-rewriting Linear Context-Free Rewriting System (LCFRS), while using a Probabilistic Discontinuous Tree Substitution Grammar (PDTSG) to improve disambiguation performance. Another system encodes the discontinuities in the labels of phrase structure trees, allowing for efficient context-free grammar parsing. The two systems demonstrate that tree fragments as used in tree-substitution grammar improve disambiguation performance while capturing non-local relations on an as-needed basis. Additionally, we present results of models that produce function tags, resulting in a more linguistically adequate model of the data. We report substantial accuracy improvements in discontinuous parsing for German, English, and Dutch, including results on spoken Dutch.