Mining a Comparable Text Corpus for a Vietnamese-French Statistical Machine Translation System

Mining a Comparable Text Corpus for a Vietnamese-French Statistical Machine Translation System
复制标题

为越南法统计机器翻译系统挖掘可比文本语料库

DOI:
10.3115/1626431.1626466
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
E. Castelli
E. Castelli
中科院分区:
--
文献类型:
--
作者:
T. Do;V. Le;B. Bigi;L. Besacier;E. Castelli

文献摘要

被引文献

相似文献

本文介绍了我们构建越南-法国统计机器翻译系统的首次尝试。由于越南语是一种资源贫乏的语言,我们专注于建立一个大型的越南语-法语平行语料库。提出了一种基于发表日期、特殊词和句子对齐结果的文档对齐方法。本文还提出了将获得的平行语料库应用于越南语-法语统计机器翻译系统的构建,其中讨论了越南语不同单位(音节、单词或其组合)的使用。
This paper presents our first attempt at constructing a Vietnamese-French statistical machine translation system. Since Vietnamese is an under-resourced language, we concentrate on building a large Vietnamese-French parallel corpus. A document alignment method based on publication date, special words and sentence alignment result is proposed. The paper also presents an application of the obtained parallel corpus to the construction of a Vietnamese-French statistical machine translation system, where the use of different units for Vietnamese (syllables, words, or their combinations) is discussed.