Paraphrasing Treebanks for Stochastic Realization Ranking

Paraphrasing Treebanks for Stochastic Realization Ranking
复制标题

解释随机实现排名的树库

DOI:
--
复制
发表时间:
2004
期刊:
影响因子:
--
通讯作者:
D. Flickinger
D. Flickinger
中科院分区:
--
文献类型:
--
作者:
Erik Velldal;S. Oepen;D. Flickinger

文献摘要

被引文献

相似文献

本文介绍了一种新的方法来实现排名的任务,即选择一个给定的输入语义之间的竞争释义,产生的生成系统。我们还介绍了对称树库的概念,我们定义为(a)一组配对的表面形式和相关的语义加上(B)的表面形式和替代实现的语义集的替代分析的组合。对于列入交替分析和实现的对称树库,我们建议,使基本的语言理论明确和操作,即在一个广泛覆盖的计算语法的形式。在Redwoods(Oepen et al. [13])范式中扩展早期基于语法的树库工作,我们提出了一个完全自动化的过程来从现有资源中生成对称树库。为了评估效用的初始(虽然小),这样的“扩大”树库,我们报告的实验结果训练随机判别模型的实现排名任务。我们的工作是在挪威语-英语机器翻译项目(LOGON; Oepen等人[11])的背景下进行的。LOGON系统建立在一个相对传统的语义转换架构上-基于最小递归语义(MRS; Copestake et al. [5])-并且通常旨在将“深层”语言主干联合收割机与随机过程相结合,以进行歧义管理并提高鲁棒性。在本文中,我们专注于孤立的子任务排名的目标语言生成器的输出。对于目标语言实现,LOGON使用了LinGO英语资源语法(ERG; Flickinger [6])和LKB生成器,这是一个词汇驱动的图表生成器,接受MRS风格的输入语义(卡罗尔等人。在代表性的LOGON数据集上,生成器已经为每个输入MRS平均生成45个英语实现;请参阅图1中的示例。因为我们预计
This paper1 describes a novel approach to the task of realization ranking, i.e. the choice among competing paraphrases for a given input semantics, as produced by a generation system. We also introduce a notion of symmetric treebanks, which we define as the combination of (a) a set of pairings of surface forms and associated semantics plus (b) the sets of alternative analyses for the surface form and sets of alternate realizations of the semantics. For inclusion of alternate analyses and realizations in the symmetric treebank, we propose to make the underlying linguistic theory explicit and operational, viz. in the form of a broad-coverage computational grammar. Extending earlier work on grammar-based treebanks in the Redwoods (Oepen et al. [13]) paradigm, we present a fully automated procedure to produce a symmetric treebank from existing resources. To evaluate the utility of an initial (albeit smallish) such ‘expanded’ treebank, we report on experimental results for training stochastic discriminative models for the realization ranking task. Our work is set within the context of a Norwegian–English machine translation project (LOGON; Oepen et al. [11]). The LOGON system builds on a relatively conventional semantic transfer architecture—based on Minimal Recursion Semantics (MRS; Copestake et al. [5])—and quite generally aims to combine a ‘deep’ linguistic backbone with stochastic processes for ambiguity management and improved robustness. In this paper we focus on the isolated subtask of ranking the output of the target language generator. For target language realization, LOGON uses the LinGO English Resource Grammar (ERG; Flickinger [6]) and LKB generator, a lexically-driven chart generator that accepts MRS-style input semantics (Carroll et al. [2]). Over a representative LOGON data set, the generator already produces an average of 45 English realizations per input MRS; see Figure 1 for an example. As we expect to move to