Cleaning the Europarl Corpus for Linguistic Applications

Cleaning the Europarl Corpus for Linguistic Applications
复制标题

清理 Europarl 语料库以进行语言应用

DOI:
--
复制
发表时间:
2014
期刊:
Conference on Natural Language Processing
影响因子:
--
通讯作者:
M. Volk
M. Volk
中科院分区:
--
文献类型:
--
作者:
Johannes Graën;D. Batinic;M. Volk

文献摘要

被引文献

相似文献

我们在当前版本的欧洲议会语料库中发现了一些反复出现的错误,这些错误既源于欧洲议会的网站,也源于基于该网站的语料库编纂。最常见的错误是元数据提取不完整,导致语料库文件的文本部分存在非文本片段。平均来说,每两次发言人转换就会出现这种情况。 我们不仅通过纠正多种错误清理了欧洲议会语料库,还对所有可用语言的发言人发言内容进行了对齐,并将所有内容编译成一个新的XML结构的语料库。这有助于更精细地选择数据,例如,查询特定政治团体的发言人的演讲语料,或特定语言组合的演讲语料。
We discovered several recurring errors in the current version of the Europarl Corpus originating both from the web site of the European Parliament and the corpus compilation based thereon. The most frequent error was incompletely extracted metadata leaving non-textual fragments within the textual parts of the corpus files. This is, on average, the case for every second speaker change. We not only cleaned the Europarl Corpus by correcting several kinds of errors, but also aligned the speakers’ contributions of all available languages and compiled every- thing into a new XML-structured corpus. This facilitates a more sophisticated selection of data, e.g. querying the corpus for speeches by speakers of a particular polit- ical group or in particular language combinations.