Building a parallel corpus for monologues with clause alignment
Building a parallel corpus for monologues with clause alignment
复制标题
为具有子句对齐的独白构建平行语料库
DOI:
--
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
Hideki Tanaka
中科院分区:
文献类型:
--
作者:
H. Kashioka;Takehiko Maruyama;Hideki Tanaka
Many studies have been reported in the domain of speech-to-speech machine translation systems for travel conversation use. Therefore, a large number of travel domain corpora have become available in recent years. From a wider viewpoint, speech-to-speech systems are required for many purposes other than travel conversation. One of these is monologues (e.g., TV news, lectures, technical presentations). However, in monologues, sentences tend to be long and complicated, which often causes problems for parsing and translation. Therefore, we need a suitable translation unit, rather than the sentence. We propose the clause as a unit for translation. To develop a speech-to-speech machine translation system for monologues based on the clause as the translation unit, we need a monologue parallel corpus with clause alignment. In this paper, we describe how to build a Japanese-English monologue parallel corpus with clauses aligned, and discuss the features of this corpus.