N-gram Counts and Language Models from the Common Crawl

N-gram Counts and Language Models from the Common Crawl
复制标题

DOI:
--
复制
发表时间:
2014-05
期刊:
--
影响因子:
--
通讯作者:
C. Buck;Kenneth Heafield;B. V. Ooyen
C. Buck;Kenneth Heafield;B. V. Ooyen
中科院分区:
其他
文献类型:
--
作者:
C. Buck;Kenneth Heafield;B. V. Ooyen

文献摘要

被引文献

相似文献

我们贡献了在Common Crawl语料库上训练的5克计数和语言模型,该语料库收集了超过90亿个网页。此版本在两个关键方面改进了Google n-gram计数:包含低计数条目和重复数据删除以减少样板文件。通过保留单例,我们能够使用Kneser-Ney平滑来构建大型语言模型。本文介绍了如何处理的语料库中出现的问题,在这种规模的数据工作的重点。我们的未修剪Kneser-Ney英语$5$-gram语言模型,建立在9750亿个去重复标记的基础上,包含超过5000亿个唯一的n-gram。通过使用大型语言模型翻译成各种语言,我们显示出0.5-1.4 BLEU的增益。
We contribute 5-gram counts and language models trained on the Common Crawl corpus, a collection over 9 billion web pages. This release improves upon the Google n-gram counts in two key ways: the inclusion of low-count entries and deduplication to reduce boilerplate. By preserving singletons, we were able to use Kneser-Ney smoothing to build large language models. This paper describes how the corpus was processed with emphasis on the problems that arise in working with data at this scale. Our unpruned Kneser-Ney English $5$-gram language model, built on 975 billion deduplicated tokens, contains over 500 billion unique n-grams. We show gains of 0.5-1.4 BLEU by using large language models to translate into various languages.