Table Union Search on Open Data

Table Union Search on Open Data
复制标题

开放数据上的表并集搜索

DOI:
--
复制
发表时间:
2018
影响因子:
2.5
通讯作者:
Renée J. Miller
Renée J. Miller
中科院分区:
计算机科学2区
文献类型:
--
作者:
F. Nargesian;Erkang Zhu;K. Pu;Renée J. Miller

文献摘要

被引文献

相似文献

我们定义了表联盟搜索问题,并提出了一个概率的解决方案,找到表是unionable与查询表在海量存储库。如果两个表共享来自同一域的属性,则它们是可联合的。我们的解决方案形式化了三个统计模型,描述了如何从集合域,语义域从本体的值,和自然语言域生成unionable属性。我们提出了一种数据驱动的方法,自动确定最佳模型用于每对属性。通过分布感知算法,我们能够找到两个表中可以联合的属性的最佳数量。为了评估准确性,我们创建并开源了开放数据表的基准。我们表明,我们的表联盟搜索在速度和准确性上优于现有的算法,用于查找相关的表和规模,以提供有效的搜索包含超过一百万个属性的开放数据存储库。
We define the table union search problem and present a probabilistic solution for finding tables that are unionable with a query table within massive repositories. Two tables are unionable if they share attributes from the same domain. Our solution formalizes three statistical models that describe how unionable attributes are generated from set domains, semantic domains with values from an ontology, and natural language domains. We propose a data-driven approach that automatically determines the best model to use for each pair of attributes. Through a distribution-aware algorithm, we are able to find the optimal number of attributes in two tables that can be unioned. To evaluate accuracy, we created and open-sourced a benchmark of Open Data tables. We show that our table union search outperforms in speed and accuracy existing algorithms for finding related tables and scales to provide efficient search over Open Data repositories containing more than one million attributes.