The brWaC Corpus: A New Open Resource for Brazilian Portuguese

The brWaC Corpus: A New Open Resource for Brazilian Portuguese
复制标题

brWaC 语料库:巴西葡萄牙语的新开放资源

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Aline Villavicencio
Aline Villavicencio
中科院分区:
--
文献类型:
--
作者:
Jorge Wagner;Rodrigo Wilkens;M. Idiart;Aline Villavicencio

文献摘要

被引文献

相似文献

在这项工作中,我们介绍了一个大型的巴西葡萄牙语网络语料库的构建过程,旨在达到与其他语言的最新水平相当的规模。我们还讨论了我们更新的句子级方法,以严格删除重复内容。按照流水线方法,6000多万个页面被爬行和过滤,其中350万个页面被选中。得到的多领域语料库brWaC由27亿个词条组成,并使用标注和句法分析信息进行了标注。在其他网络语料库中,表明内容重复的非唯一长句的发生率达到了9%,下降到只有0.5%。域名多样性也得到了最大程度的提高,12万个不同的网站提供了内容。我们正在向研究界免费提供我们的新资源,供查询和下载,以期帮助巴西葡萄牙语的处理取得新的进展。
In this work, we present the construction process of a large Web corpus for Brazilian Portuguese, aiming to achieve a size comparable to the state of the art in other languages. We also discuss our updated sentence-level approach for the strict removal of duplicated content. Following the pipeline methodology, more than 60 million pages were crawled and filtered, with 3.5 million being selected. The obtained multi-domain corpus, named brWaC, is composed by 2.7 billion tokens, and has been annotated with tagging and parsing information. The incidence of non-unique long sentences, an indication of replicated content, which reaches 9% in other Web corpora, was reduced to only 0.5%. Domain diversity was also maximized, with 120,000 different websites contributing content. We are making our new resource freely available for the research community, both for querying and downloading, in the expectation of aiding in new advances for the processing of Brazilian Portuguese.