A massively parallel corpus: the Bible in 100 languages.

A massively parallel corpus: the Bible in 100 languages.
复制标题

DOI:
10.1007/s10579-014-9287-y
复制
发表时间:
2015
影响因子:
2.7
通讯作者:
Steedman M
Steedman M
中科院分区:
计算机科学4区
文献类型:
--
作者:
Christodouloupoulos C;Steedman M

文献摘要

被引文献

相似文献

我们描述了一个基于100个圣经译本的大规模平行语料库的创建。我们讨论了获取和处理原始材料的一些困难,以及圣经作为自然语言处理语料库的潜力。最后对收集到的语料库进行了统计分析,并与其他英语语料库进行了详细的比较。
We describe the creation of a massively parallel corpus based on 100 translations of the Bible. We discuss some of the difficulties in acquiring and processing the raw material as well as the potential of the Bible as a corpus for natural language processing. Finally we present a statistical analysis of the corpora collected and a detailed comparison between the English translation and other English corpora.