Mining a Comparable Text Corpus for a Vietnamese-French Statistical Machine Translation System
Mining a Comparable Text Corpus for a Vietnamese-French Statistical Machine Translation System
复制标题
为越南法统计机器翻译系统挖掘可比文本语料库
DOI:
10.3115/1626431.1626466
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
E. Castelli
中科院分区:
文献类型:
--
作者:
T. Do;V. Le;B. Bigi;L. Besacier;E. Castelli
This paper presents our first attempt at constructing a Vietnamese-French statistical machine translation system. Since Vietnamese is an under-resourced language, we concentrate on building a large Vietnamese-French parallel corpus. A document alignment method based on publication date, special words and sentence alignment result is proposed. The paper also presents an application of the obtained parallel corpus to the construction of a Vietnamese-French statistical machine translation system, where the use of different units for Vietnamese (syllables, words, or their combinations) is discussed.