Using the Europarl corpus for cross-linguistic research

Using the Europarl corpus for cross-linguistic research
复制标题

使用 Europarl 语料库进行跨语言研究

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
T. Meyer
T. Meyer
中科院分区:
--
文献类型:
--
作者:
Bruno Cartoni;S. Zufferey;T. Meyer

文献摘要

被引文献

相似文献

Europarl是一个包含欧洲议会辩论记录的大型多语种语料库。本文提出了一种从Europarl中抽取不同语料库的方法:单语和多语可比语料库,以及平行语料库。使用最先进的同质性衡量标准,我们表明这些语料库非常相似。此外,我们认为它们为语言学和翻译研究的各个领域的研究提供了许多优势,并讨论了它们的一些局限性。最后,我们回顾了一些以前利用这些语料库的研究,强调了Europarl提供的可能性。
Europarl is a large multilingual corpus containing the minutes of the debates at the European Parliament. This article presents a method to extract different corpora from Europarl: monolingual and multilingual comparable corpora, as well as parallel corpora. Using state-of-the-art measures of homogeneity, we show that these corpora are very similar. In addition, we argue that they present many advantages for research in various fields of linguistics and translation studies, and we also discuss some of their limitations. We conclude by reviewing a number of previous studies that made use of these corpora, emphasizing in each case the possibilities offered by Europarl.