The NAIST Simultaneous Translation Corpus

The NAIST Simultaneous Translation Corpus
复制标题

NAIST 同声翻译语料库

DOI:
10.1007/978-981-10-6199-8_11
复制
发表时间:
2018
期刊:
影响因子:
5.8
通讯作者:
T. Toda
T. Toda
中科院分区:
生物学2区
文献类型:
--
作者:
Graham Neubig;Hiroaki Shimizu;S. Sakti;Satoshi Nakamura;T. Toda

文献摘要

被引文献

相似文献

本章介绍了在奈良科学技术学院(NAIST)收集的英语/日本/日语同时解释语料库。语料库的两个主要特征将其与众不同。首先是它包含具有不同经验的专业同时口译员的记录解释结果。这使得可以比较不同级别的口译员之间的差异,从而阐明了口译员经验对结果的客观和主观品质的影响。第二个功能是该语料库的一部分也已翻译。当在没有时间限制的情况下(使用翻译数据)或与时间限制(使用同时的解释数据),从文本翻译特定的谈话时,可以比较和对比结果。该语料库总共包含387,000个单词的数据,其中涵盖了讲座和新闻的材料。所有转录均为时间对齐。该语料库将有助于分析解释方式的差异,也可以用作同时解释系统构建的参考。
This chapter describes an English-Japanese/Japanese-English simultaneous interpretation corpus collected at the Nara Institute of Science and Technology (NAIST). There are two main features of the corpus that set it apart from others. The first is that it contains recorded interpretation results from professional simultaneous interpreters with different amounts of experience. This makes it possible to compare the differences between interpreters of different levels, elucidating the effect of interpreter experience on the objective and subjective qualities of results. The second feature is that part of the corpus also has been translated. This data makes it possible to compare and contrast the results when a particular talk is translated from text without time constraints (using the translation data) or from speech with time constraints (using the simultaneous interpretation data). The corpus contains a total of 387k words worth of data, with the material covering lectures and news. All transcriptions are time aligned. The corpus will be helpful to analyze differences in interpretation styles, and may also be used as a reference in the construction of simultaneous interpretation systems.