† National Institute of Information and Communications Technology ‡ ATR Spoken Language Communication Research Labs Hikaridai 2-2-2, Keihanna Science City, 619-0288 Kyoto {Michael.Paul,Eiichiro.Sumita}@{nict.go.jp,atr.j p}
† National Institute of Information and Communications Technology ‡ ATR Spoken Language Communication Research Labs Hikaridai 2-2-2, Keihanna Science City, 619-0288 Kyoto {Michael.Paul,Eiichiro.Sumita}@{nict.go.jp,atr.j p}
复制标题
† 国立信息通信技术研究所 ‡ ATR 口语交流研究实验室 Hikaridai 2-2-2, Keihanna Science City, 619-0288 京都 {Michael.Paul,Eiichiro.Sumita}@{nict.go.jp,atr.j p }
DOI:
--
复制
发表时间:
2006
期刊:
影响因子:
--
通讯作者:
E. Sumita
中科院分区:
文献类型:
--
作者:
Michael Paul;E. Sumita
This paper proposes the usage of variant corpora, i.e., parallel text corpora that are equal in meaning but use different ways to express content, in order to improve corpus-based machine translation. The usage of multiple training corpora of the same content with different sources results in variant models that focus on specific linguistic phenomena covered by the respective corpus. The proposed method applies each variant model separately resulting in multiple translation hypotheses which are selectively combined according to statistical models. The proposed method outperforms the conventional approach of merging all variants by reducing translation ambiguities and exploiting the strengths of each variant model.