A description of the English-Norwegian parallel corpus : Compilation and further developments

A description of the English-Norwegian parallel corpus : Compilation and further developments
复制标题

英语-挪威语平行语料库的描述:编译和进一步发展

DOI:
10.1075/ijcl.4.2.01oks
复制
发表时间:
1999
影响因子:
1
通讯作者:
Signe Oksefjell
Signe Oksefjell
中科院分区:
人文科学4区
文献类型:
--
作者:
Signe Oksefjell

文献摘要

被引文献

相似文献

本文介绍了英语-挪威语平行语料库(ENPC)的编制过程中的几个重要步骤。ENPC包含50个英语原文摘录及其挪威语译文和50个挪威原文摘录及其英语译文,共约260万字。即使这一过程中最耗时的部分是为语料库准备文本摘录,但也有很大一部分重点放在软件的开发上,特别是处理平行文本的浏览器和将同一文本的原文和译文连接起来的对齐程序。文本本身的准备包括扫描、校对、标记和对齐。虽然ENPC已经完成,但ENPC项目仍在发展中,本文将提到最新的扩展,例如添加更多的语言,编译同一文本的多个翻译(同一语言),词性标记,以及标记ENPC中的直接语音和思想。
This paper gives an introduction to the most important steps in the process of compiling the English-Norwegian Parallel Corpus (ENPC), which contains 50 original English text extracts with their translations into Norwegian and 50 original Norwegian text extracts with their translations into English, in all about 2.6 million words. Even if the most time-consuming part of the process is to prepare the text extracts for the corpus, much of the focus has also been on the development of software, notably a browser handling parallel texts and an alignment program linking the original and translated versions of the same text. The preparation of the texts themselves includes scanning, proofreading, mark-up, and alignment. Although the ENPC is completed, the ENPC project is still developing, and the most recent extensions will be mentioned in this paper, such as adding more languages, compiling multiple translations (in the same language) of the same text, part-of-speech-tagging, and marking direct speech and thought in the ENPC.