A Corpus-based Probabilistic Grammar with Only Two Non-terminals

A Corpus-based Probabilistic Grammar with Only Two Non-terminals
复制标题

一种基于语料库的只有两个非终结符的概率语法

DOI:
--
复制
发表时间:
1995
期刊:
International Workshop/Conference on Parsing Technologies
影响因子:
--
通讯作者:
R. Grishman
R. Grishman
中科院分区:
--
文献类型:
--
作者:
S. Sekine;R. Grishman

文献摘要

被引文献

相似文献

像Penn Tree Bank这样的大型语法括号语料库的可用性为我们提供了自动构建或训练广泛覆盖语法的机会,特别是训练概率语法。最近的一些解析实验也表明,在选择正确的解析时,生成概率依赖于上下文的语法比上下文无关的语法更有效。为了最大限度地利用上下文,我们从Penn Tree Bank版本2中自动构建了一个语法,其中符号S和NP是唯一真正的非终结符,而其他非终结符或语法节点实际上嵌入到S和NP规则的右侧。例如,从树库中提取的规则之一是S -> NP VBX JJ CC VBX NP[1](其中NP是非终结符,其他符号是终结符——树库的词性标签)。与此扩展相关的树库中最常见的结构是(S NP (VP (VP VBX (ADJ JJ) CC (VP VBX NP))))[2])。因此,如果我们的解析器在解析句子时使用规则[1],它将为句子的相应部分生成结构[2]。使用94%的Penn Tree Bank进行训练,我们提取了32,296条不同的规则(23,386条用于S, 8,910条用于NP)。我们还基于高频模式构建了一个较小的语法版本,以便在较大的语法由于内存限制而无法生成解析时用作备份。我们将这个解析器应用于1989个华尔街日报句子(与训练集分开,没有句子长度限制)。在解析的1,899个句子中,无交叉句的比例为33.9%,Parseval查全率和查准率分别为73.43%和72.61%。
The availability of large, syntactically-bracketed corpora such as the Penn Tree Bank affords us the opportunity to automatically build or train broad-coverage grammars, and in particular to train probabilistic grammars. A number of recent parsing experiments have also indicated that grammars whose production probabilities are dependent on the context can be more effective than context-free grammars in selecting a correct parse. To make maximal use of context, we have automatically constructed, from the Penn Tree Bank version 2, a grammar in which the symbols S and NP are the only real nonterminals, and the other non-terminals or grammatical nodes are in effect embedded into the right-hand-sides of the S and NP rules. For example, one of the rules extracted from the tree bank would be S -> NP VBX JJ CC VBX NP [1] ( where NP is a non-terminal and the other symbols are terminals – part-of-speech tags of the Tree Bank). The most common structure in the Tree Bank associated with this expansion is (S NP (VP (VP VBX (ADJ JJ) CC (VP VBX NP)))) [2]. So if our parser uses rule [1] in parsing a sentence, it will generate structure [2] for the corresponding part of the sentence. Using 94% of the Penn Tree Bank for training, we extracted 32,296 distinct rules ( 23,386 for S, and 8,910 for NP). We also built a smaller version of the grammar based on higher frequency patterns for use as a back-up when the larger grammar is unable to produce a parse due to memory limitation. We applied this parser to 1,989 Wall Street Journal sentences (separate from the training set and with no limit on sentence length). Of the parsed sentences (1,899), the percentage of no-crossing sentences is 33.9%, and Parseval recall and precision are 73.43% and 72 .61%.