Towards History-based Grammars: Using Richer Models for Probabilistic Parsing

Towards History-based Grammars: Using Richer Models for Probabilistic Parsing
复制标题

迈向基于历史的语法:使用更丰富的模型进行概率解析

DOI:
10.3115/981574.981579
复制
发表时间:
1993
期刊:
--
影响因子:
--
通讯作者:
S. Roukos
S. Roukos
中科院分区:
--
文献类型:
--
作者:
Ezra Black;F. Jelinek;J. Lafferty;David M. Magerman;R. Mercer;S. Roukos

文献摘要

被引文献

相似文献

我们描述了一个自然语言的生成概率模型,我们称之为HBG,它利用详细的语言信息来解决歧义。HBG以一种新颖的方式将来自解析树的词汇、句法、语义和结构信息整合到消歧过程中。我们使用括号句子的语料库,称为树库,结合决策树构建来梳理出解析树的相关方面,这将确定句子的正确解析。这与通过通常的语言内省来进一步裁剪语法以生成正确的解析的通常方法形成对比。在对现有最好的鲁棒概率解析模型之一(我们称之为P-CFG)的头对头测试中,HBG模型明显优于P-CFG,将解析准确率从60%提高到75%,错误减少了37%。
We describe a generative probabilistic model of natural language, which we call HBG, that takes advantage of detailed linguistic information to resolve ambiguity. HBG incorporates lexical, syntactic, semantic, and structural information from the parse tree into the disambiguation process in a novel way. We use a corpus of bracketed sentences, called a Treebank, in combination with decision tree building to tease out the relevant aspects of a parse tree that will determine the correct parse of a sentence. This stands in contrast to the usual approach of further grammar tailoring via the usual linguistic introspection in the hope of generating the correct parse. In head-to-head tests against one of the best existing robust probabilistic parsing models, which we call P-CFG, the HBG model significantly outperforms P-CFG, increasing the parsing accuracy rate from 60% to 75%, a 37% reduction in error.