Building a parallel corpus for monologues with clause alignment

Building a parallel corpus for monologues with clause alignment
复制标题

为具有子句对齐的独白构建平行语料库

DOI:
--
复制
发表时间:
2003
期刊:
Machine Translation Summit
影响因子:
--
通讯作者:
Hideki Tanaka
Hideki Tanaka
中科院分区:
--
文献类型:
--
作者:
H. Kashioka;Takehiko Maruyama;Hideki Tanaka

文献摘要

被引文献

相似文献

许多研究已经报道了在语音到语音机器翻译系统的旅游会话使用的领域。因此,大量的旅游领域语料库已成为近年来。从更广泛的角度来看,语音到语音系统需要用于许多目的,而不是旅行对话。其中之一是独白(例如,电视新闻、讲座、技术演示)。然而,在独白中,句子往往是长的和复杂的,这往往会导致分析和翻译的问题。因此,我们需要一个合适的翻译单位,而不是句子。我们建议将该条款作为翻译单位。为了开发一个以小句为翻译单位的独白语到语机器翻译系统,我们需要一个具有小句对齐功能的独白语平行语料库。本文描述了如何建立一个具有小句对齐的日英独白平行语料库,并讨论了该语料库的特点。
Many studies have been reported in the domain of speech-to-speech machine translation systems for travel conversation use. Therefore, a large number of travel domain corpora have become available in recent years. From a wider viewpoint, speech-to-speech systems are required for many purposes other than travel conversation. One of these is monologues (e.g., TV news, lectures, technical presentations). However, in monologues, sentences tend to be long and complicated, which often causes problems for parsing and translation. Therefore, we need a suitable translation unit, rather than the sentence. We propose the clause as a unit for translation. To develop a speech-to-speech machine translation system for monologues based on the clause as the translation unit, we need a monologue parallel corpus with clause alignment. In this paper, we describe how to build a Japanese-English monologue parallel corpus with clauses aligned, and discuss the features of this corpus.