WTR: A Test Collection for Web Table Retrieval

WTR: A Test Collection for Web Table Retrieval
复制标题

DOI:
10.1145/3404835.3463260
复制
发表时间:
2021-05
期刊:
Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Zhiyu Chen;Shuo Zhang;B. Davison
Zhiyu Chen;Shuo Zhang;B. Davison
中科院分区:
其他
文献类型:
--
作者:
Zhiyu Chen;Shuo Zhang;B. Davison

文献摘要

相似文献

我们描述了用于Web表检索任务的测试集的开发、特征和可用性,该测试集使用从Common Crawl中提取的大规模Web表语料库。由于Web表通常具有丰富的上下文信息,例如页面标题和周围段落,因此我们不仅提供查询表对的相关性判断,还提供查询表上下文对与查询的相关性判断,而这些判断被以前的测试集忽略。为了促进未来使用该基准的研究,我们提供了有关如何预处理数据集的详细信息,以及传统和最近提出的表检索方法的基线结果。我们的实验结果表明,正确使用上下文标签可以使以前的表检索方法受益。
We describe the development, characteristics and availability of a test collection for the task of Web table retrieval, which uses a large-scale Web Table Corpora extracted from the Common Crawl. Since a Web table usually has rich context information such as the page title and surrounding paragraphs, we not only provide relevance judgments of query-table pairs, but also the relevance judgments of query-table context pairs with respect to a query, which are ignored by previous test collections. To facilitate future research with this benchmark, we provide details about how the dataset is pre-processed and also baseline results from both traditional and recently proposed table retrieval methods. Our experimental results show that proper usage of context labels can benefit previous table retrieval methods.