ClueWeb22: 10 Billion Web Documents with Rich Information

ClueWeb22: 10 Billion Web Documents with Rich Information
复制标题

DOI:
10.1145/3477495.3536321
复制
发表时间:
2022-07
期刊:
Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Arnold Overwijk;Chenyan Xiong;Jamie Callan
Arnold Overwijk;Chenyan Xiong;Jamie Callan
中科院分区:
其他
文献类型:
--
作者:
Arnold Overwijk;Chenyan Xiong;Jamie Callan

文献摘要

相似文献

ClueWeb22 是 ClueWeb 数据集系列的最新版本,是行业和学术界一年多合作的成果。其设计受到学术界的研究需求和大规模工业系统的现实需求的影响。与早期的 ClueWeb 数据集相比,ClueWeb22 语料库规模更大、种类更多、文档质量更高。它的核心是原始 HTML,但它包含文档的干净文本版本,以降低进入门槛。 ClueWeb22 的多个方面首次以这种规模提供给研究社区,例如,渲染网页的视觉表示、从 HTML 文档解析的结构化信息以及文档分布(域、语言和主题)与商业网络搜索的对齐。本次演讲分享了 ClueWeb22 的设计和构建,并讨论了它的新功能。我们相信这个更新、更大、更丰富的 ClueWeb 语料库将实现并支持 IR、NLP 和深度学习领域的广泛研究。
ClueWeb22, the newest iteration of the ClueWeb line of datasets, is the result of more than a year of collaboration between industry and academia. Its design is influenced by the research needs of the academic community and the real-world needs of large-scale industry systems. Compared with earlier ClueWeb datasets, the ClueWeb22 corpus is larger, more varied, and has higher-quality documents. Its core is raw HTML, but it includes clean text versions of documents to lower the barrier to entry. Several aspects of ClueWeb22 are available to the research community for the first time at this scale, for example, visual representations of rendered web pages, parsed structured information from the HTML document, and the alignment of document distributions (domains, languages, and topics) to commercial web search. This talk shares the design and construction of ClueWeb22, and discusses its new features. We believe this newer, larger, and richer ClueWeb corpus will enable and support a broad range of research in IR, NLP, and deep learning.