Paraphrasing Treebanks for Stochastic Realization Ranking
Paraphrasing Treebanks for Stochastic Realization Ranking
复制标题
解释随机实现排名的树库
DOI:
--
复制
发表时间:
2004
期刊:
影响因子:
--
通讯作者:
D. Flickinger
中科院分区:
文献类型:
--
作者:
Erik Velldal;S. Oepen;D. Flickinger
This paper1 describes a novel approach to the task of realization ranking, i.e. the choice among competing paraphrases for a given input semantics, as produced by a generation system. We also introduce a notion of symmetric treebanks, which we define as the combination of (a) a set of pairings of surface forms and associated semantics plus (b) the sets of alternative analyses for the surface form and sets of alternate realizations of the semantics. For inclusion of alternate analyses and realizations in the symmetric treebank, we propose to make the underlying linguistic theory explicit and operational, viz. in the form of a broad-coverage computational grammar. Extending earlier work on grammar-based treebanks in the Redwoods (Oepen et al. [13]) paradigm, we present a fully automated procedure to produce a symmetric treebank from existing resources. To evaluate the utility of an initial (albeit smallish) such ‘expanded’ treebank, we report on experimental results for training stochastic discriminative models for the realization ranking task. Our work is set within the context of a Norwegian–English machine translation project (LOGON; Oepen et al. [11]). The LOGON system builds on a relatively conventional semantic transfer architecture—based on Minimal Recursion Semantics (MRS; Copestake et al. [5])—and quite generally aims to combine a ‘deep’ linguistic backbone with stochastic processes for ambiguity management and improved robustness. In this paper we focus on the isolated subtask of ranking the output of the target language generator. For target language realization, LOGON uses the LinGO English Resource Grammar (ERG; Flickinger [6]) and LKB generator, a lexically-driven chart generator that accepts MRS-style input semantics (Carroll et al. [2]). Over a representative LOGON data set, the generator already produces an average of 45 English realizations per input MRS; see Figure 1 for an example. As we expect to move to