Challenges for Dataset Search

Challenges for Dataset Search
复制标题

数据集搜索的挑战

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Database Systems for Advanced Applications
影响因子:
--
通讯作者:
Kristin Tufte
Kristin Tufte
中科院分区:
--
文献类型:
--
作者:
D. Maier;V. M. Megler;Kristin Tufte

文献摘要

被引文献

相似文献

随着共享科学档案的规模和种类不断增加,数据集的排名搜索已成为一种需求。我们自己的研究表明,IR 风格、基于特征的相关性评分可以成为科学档案中数据发现的有效工具。然而,随着档案规模的扩大,保持交互式响应时间将是一个挑战。我们在此报告我们对数据集搜索服务 Data Near Here 的性能技术的探索。我们提供了评估系统中滤波器重启技术的结果样本,包括两种变体:自适应松弛和收缩。然后我们概述了该领域的进一步研究方向。
Ranked search of datasets has emerged as a need as shared scientific archives grow in size and variety. Our own have shown that IR-style, feature-based relevance scoring can be an effective tool for data discovery in scientific archives. However, maintaining interactive response times as archives scale will be a challenge. We report here on our exploration of performance techniques for Data Near Here, a dataset search service. We present a sample of results evaluating filter-restart techniques in our system, including two variations, adaptive relaxation and contraction. We then outline further directions for research in this domain.