Using the Europarl corpus for cross-linguistic research
Using the Europarl corpus for cross-linguistic research
复制标题
使用 Europarl 语料库进行跨语言研究
DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
T. Meyer
中科院分区:
文献类型:
--
作者:
Bruno Cartoni;S. Zufferey;T. Meyer
Europarl is a large multilingual corpus containing the minutes of the debates at the European Parliament. This article presents a method to extract different corpora from Europarl: monolingual and multilingual comparable corpora, as well as parallel corpora. Using state-of-the-art measures of homogeneity, we show that these corpora are very similar. In addition, we argue that they present many advantages for research in various fields of linguistics and translation studies, and we also discuss some of their limitations. We conclude by reviewing a number of previous studies that made use of these corpora, emphasizing in each case the possibilities offered by Europarl.