Source reordering using MaxEnt classifiers and supertags

Source reordering using MaxEnt classifiers and supertags
复制标题

DOI:
--
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
Maxim Khalilov;K. Sima'an
Maxim Khalilov;K. Sima'an
中科院分区:
其他
文献类型:
--
作者:
Maxim Khalilov;K. Sima'an

文献摘要

被引文献

相似文献

源语言重新排序可以被看作是以这样的方式排列源词的顺序的预处理任务,即所产生的排列允许尽可能单调的翻译过程。我们探索了一种简单但有效的源重排序算法,它作为源串转换的级联工作,每个转换包括交换单个相邻单词对的位置,以便展开候选交叉对齐。将一对单词互换的决定建模为一个二进制分类任务,该任务被表示为对数线性模型,并在最大熵(MaxEnt)下进行训练。我们试验了由这两个单词的局部邻域以及称为超标签的词汇句法表示组成的特征。我们在英语到荷兰语EuroParl翻译任务上的实验表明,级联对齐展开略微改善了使用基于距离和面向词汇化块重新排序的最新短语翻译系统的性能。
Source language reordering can be seen as the preprocessing task of permuting the order of the source words in such a way that the resulting permutation allows as monotone a translation process as possible. We explore a simple but effective source reordering algorithm that works as a cascade of source string transforms, each consisting of swapping the positions of a single pair of adjacent words in order to unfold a candidate pair of crossing alignments. The decision to swap a pair of words is modelled as a binary classification task formulated as a log-linear model and trained under maximum entropy (MaxEnt). We experiment with features that consist of the local neighborhood of both words as well as lexico-syntactic representations known as supertags. Our experiments on the English-to-Dutch EuroParl translation task show that the cascaded alignment unfolding slightly improves the performance of a state-of-the-art phrase translation system that uses distance-based and lexicalized block-oriented reordering.