SANTOS: Relationship-based Semantic Table Union Search

SANTOS: Relationship-based Semantic Table Union Search
复制标题

SANTOS:基于关系的语义表联合搜索

DOI:
10.1145/3588689
复制
发表时间:
2023
期刊:
Proceedings of the ACM on Management of Data
影响因子:
--
通讯作者:
Riedewald, Mirek
Riedewald, Mirek
中科院分区:
--
文献类型:
--
作者:
Khatiwada, Aamod;Fan, Grace;Shraga, Roee;Chen, Zixuan;Gatterbauer, Wolfgang;Miller, Renée J.;Riedewald, Mirek

文献摘要

参考文献

被引文献

相似文献

用于可联合表搜索的现有技术使用元数据(表必须具有相同或相似的模式)或基于列的度量(例如,表中的值应该从相同的域中提取)来定义可联合性。在这项工作中,我们介绍了使用表中的列对之间的语义关系,以提高联合搜索的准确性。因此,我们引入了一个新的概念unionability考虑列之间的关系,连同列的语义,在原则上的方式。为此,我们提出了两种新方法来发现列对之间的语义关系。第一个使用现有的知识库(KB),第二个(我们称之为“合成KB”)使用来自数据湖本身的知识。我们采用了现有的表联盟搜索基准,并提出了新的(开放)基准,代表小型和大型真实的数据湖。我们表明,我们的新unionability搜索算法,称为桑托斯,优于国家的最先进的工会搜索,使用各种各样的基于列的语义,包括词嵌入和正则表达式。我们的经验表明,我们的合成KB提高了联盟搜索的准确性,表示关系语义,可能不包含在一个可用的KB。这一结果暗示了一个充满希望的未来,从有限的知识库覆盖的数据湖创建合成知识库,并使用它们进行联合搜索。
Existing techniques for unionable table search define unionability using metadata (tables must have the same or similar schemas) or column-based metrics (for example, the values in a table should be drawn from the same domain). In this work, we introduce the use of semantic relationships between pairs of columns in a table to improve the accuracy of the union search. Consequently, we introduce a new notion of unionability that considers relationships between columns, together with the semantics of columns, in a principled way. To do so, we present two new methods to discover the semantic relationships between pairs of columns. The first uses an existing knowledge base (KB), and the second (which we call a "synthesized KB") uses knowledge from the data lake itself. We adopt an existing Table Union Search benchmark and present new (open) benchmarks that represent small and large real data lakes. We show that our new unionability search algorithm, called SANTOS, outperforms a state-of-the-art union search that uses a wide variety of column-based semantics, including word embeddings and regular expressions. We show empirically that our synthesized KB improves the accuracy of union search by representing relationship semantics that may not be contained in an available KB. This result hints at a promising future of creating synthesized KBs from data lakes with limited KB coverage and using them for union search.
ACM SIGMOD 2020 论文的再现性报告:“为数据集成任务创建异构关系数据集的嵌入”
DOI: --
发表时间: 2021
期刊:
影响因子: --
作者:
Riccardo Cappuzzo;Paolo Papotti;Saravanan Thirumuruganathan
通讯作者: Saravanan Thirumuruganathan
QuTE:回答来自 Web 表的数量查询
DOI: --
发表时间: 2021
期刊: SIGMOD Conference
影响因子: --
作者:
Vinh Thinh Ho;K. Pal;G. Weikum
通讯作者: G. Weikum
DOI: 10.1145/3318464.3389726
发表时间: 2020-06
期刊: Proceedings. ACM-SIGMOD International Conference on Management of Data
影响因子: --
作者:
Zhang Y;Ives ZG
通讯作者: Ives ZG
开放数据集成
DOI: --
发表时间: 2018
影响因子: 2.5
作者:
Renée J. Miller
通讯作者: Renée J. Miller
DOI: --
发表时间: 2016
期刊:
影响因子: --
作者:
Brent Glays
通讯作者: Brent Glays