Leveraging Schema Labels to Enhance Dataset Search

Leveraging Schema Labels to Enhance Dataset Search
复制标题

DOI:
10.1007/978-3-030-45439-5_18
复制
发表时间:
2020-03-17
期刊:
Advances in Information Retrieval
影响因子:
--
通讯作者:
Davison BD
Davison BD
中科院分区:
其他
文献类型:
--
作者:
Chen Z;Jia H;Heflin J;Davison BD

文献摘要

被引文献

相似文献

搜索引擎检索所需数据集的能力对于数据共享和重用非常重要。现有的数据集搜索引擎通常依赖于将查询匹配到数据集描述。然而,用户可能没有足够的先验知识来使用与描述文本匹配的术语来编写查询。我们提出了一种新的模式标签生成模型,该模型根据数据集表内容生成可能的模式标签。我们将生成的模式标签合并到一个混合排名模型中,该模型不仅考虑查询和数据集元数据之间的相关性,还考虑查询和生成的模式标签之间的相似性。为了在真实世界的数据集上评估我们的方法,我们专门为数据集检索任务创建了一个新的基准。实验结果表明,与基线方法相比,该方法可以有效地提高数据集检索任务的精度和NDCG分数。我们还测试了一组维基百科表,以表明从模式标签生成的功能可以提高无监督和监督的Web表检索任务。
A search engine’s ability to retrieve desirable datasets is important for data sharing and reuse. Existing dataset search engines typically rely on matching queries to dataset descriptions. However, a user may not have enough prior knowledge to write a query using terms that match with description text. We propose a novel schema label generation model which generates possible schema labels based on dataset table content. We incorporate the generated schema labels into a mixed ranking model which not only considers the relevance between the query and dataset metadata but also the similarity between the query and generated schema labels. To evaluate our method on real-world datasets, we create a new benchmark specifically for the dataset retrieval task. Experiments show that our approach can effectively improve the precision and NDCG scores of the dataset retrieval task compared with baseline methods. We also test on a collection of Wikipedia tables to show that the features generated from schema labels can improve the unsupervised and supervised web table retrieval task as well.