Overview of the Patent Translation Task at the NTCIR-7 Workshop

Overview of the Patent Translation Task at the NTCIR-7 Workshop
复制标题

DOI:
--
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
Atsushi Fujii;M. Utiyama;Mikio Yamamoto;T. Utsuro
Atsushi Fujii;M. Utiyama;Mikio Yamamoto;T. Utsuro
中科院分区:
其他
文献类型:
--
作者:
Atsushi Fujii;M. Utiyama;Mikio Yamamoto;T. Utsuro

文献摘要

被引文献

相似文献

为了帮助机器翻译的研究和开发,我们制作了日语/英语机器翻译的测试集,并在第七届NTCIR研讨会上执行了专利翻译任务。为了获得平行语料库,我们提取了在日本和美国发表的相同或相关发明的专利文件。我们的测试集包括大约200万句对日语和英语,这是自动提取我们的平行语料库。这些句子对可以用来训练和评估机器翻译系统。我们的测试集还包括跨语言专利检索的搜索主题,可用于评估机器翻译对跨语言检索专利文档的贡献。本文描述了我们的测试集,评估机器翻译的方法,以及参与我们任务的研究小组的评估结果。我们的研究是第一次利用专利信息进行机器翻译评估的重要探索。
To aid research and development in machine translation, we have produced a test collection for Japanese/English machine translation and performed the Patent Translation Task at the Seventh NTCIR Workshop. To obtain a parallel corpus, we extracted patent documents for the same or related inventions published in Japan and the United States. Our test collection includes approximately 2 000 000 sentence pairs in Japanese and English, which were extracted automatically from our parallel corpus. These sentence pairs can be used to train and evaluate machine translation systems. Our test collection also includes search topics for cross-lingual patent retrieval, which can be used to evaluate the contribution of machine translation to retrieving patent documents across languages. This paper describes our test collection, methods for evaluating machine translation, and evaluation results for research groups participated in our task. Our research is the first significant exploration into utilizing patent information for the evaluation of machine translations.