The Web Is Your Oyster - Knowledge-Intensive NLP against a Very Large Web Corpus

The Web Is Your Oyster - Knowledge-Intensive NLP against a Very Large Web Corpus
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Aleksandra Piktus;F. Petroni;Vladimir Karpukhin;Dmytro Okhonko;Samuel Broscheit;Gautier Izacard;Patrick Lewis;Barlas Ouguz;Edouard Grave;Wen-tau Yih;Sebastian Riedel
Aleksandra Piktus;F. Petroni;Vladimir Karpukhin;Dmytro Okhonko;Samuel Broscheit;Gautier Izacard;Patrick Lewis;Barlas Ouguz;Edouard Grave;Wen-tau Yih;Sebastian Riedel
中科院分区:
其他
文献类型:
--
作者:
Aleksandra Piktus;F. Petroni;Vladimir Karpukhin;Dmytro Okhonko;Samuel Broscheit;Gautier Izacard;Patrick Lewis;Barlas Ouguz;Edouard Grave;Wen-tau Yih;Sebastian Riedel

文献摘要

被引文献

相似文献

为了满足现实世界应用的不断增长的需求,知识密集型NLP(KI-NLP)的研究应通过捕捉真正开放域环境的挑战来提高:网络规模的知识,缺乏结构,质量不一致和不一致噪音。为此,我们提出了一个新的设置,用于评估现有知识密集型任务,其中我们将背景语料库推广到通用的Web快照。我们研究了依靠知识的NLP任务(事实或常识),并要求系统使用CCNET的子集 - Sphere语料库作为知识来源。与Wikipedia相反,否则在Ki-NLP中是常见的背景语料库,球体是更大的数量级,更好地反映了网络上知识的全部多样性。尽管覆盖范围,规模挑战,缺乏结构和质量较低的潜在差距,但我们发现从球体中取回可以使最先进的系统在多个任务上匹配甚至超过基于Wikipedia的模型。我们还观察到,虽然密集的指数可以胜过Wikipedia上稀疏的BM25基线,但在球面上,这是不可能的。为了促进进一步的研究并最大程度地减少社区对专有的黑盒搜索引擎的依赖,我们共享指标,评估指标和基础架构。
In order to address increasing demands of real-world applications, the research for knowledge-intensive NLP (KI-NLP) should advance by capturing the challenges of a truly open-domain environment: web-scale knowledge, lack of structure, inconsistent quality and noise. To this end, we propose a new setup for evaluating existing knowledge intensive tasks in which we generalize the background corpus to a universal web snapshot. We investigate a slate of NLP tasks which rely on knowledge - either factual or common sense, and ask systems to use a subset of CCNet - the Sphere corpus - as a knowledge source. In contrast to Wikipedia, otherwise a common background corpus in KI-NLP, Sphere is orders of magnitude larger and better reflects the full diversity of knowledge on the web. Despite potential gaps in coverage, challenges of scale, lack of structure and lower quality, we find that retrieval from Sphere enables a state of the art system to match and even outperform Wikipedia-based models on several tasks. We also observe that while a dense index can outperform a sparse BM25 baseline on Wikipedia, on Sphere this is not yet possible. To facilitate further research and minimise the community's reliance on proprietary, black-box search engines, we share our indices, evaluation metrics and infrastructure.