课题基金 / 基金详情

III: Medium: Table-as-Query: Unifying Data Discovery and Alignment

III: Medium: Table-as-Query: Unifying Data Discovery and Alignment
III:媒介:表即查询:统一数据发现和对齐
批准号:
1956096
负责人:
Renee Miller
金额:
$100.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-08-01 至 2024-07-31

项目摘要

项目成果

Renee Miller的其他基金

相似基金

相关文献

中文摘要
翻译
在信息提取的进步和重视机构开放性和透明度的社会趋势的推动下,结构化数据正在以压倒性的速度产生和共享。开放数据共享是支持机构透明度的核心,但如果无法找到共享数据,并且无法有效地与数据科学家、记者和其他人正在研究的其他数据保持一致,则无法实现透明度。该项目将从根本上为开放数据共享的新科学做出贡献。对包含结构化数据的异构表存储库上的数据发现和集成的要求与联合数据集成(例如,集成企业内的所有数据)或数据交换(其中,数据在一小部分自治对等体之间交换,例如在两个机构之间)根本不同。该项目将为表存储库中的数据发现(识别、对齐和集成表)奠定理论基础。它将有助于开发正确的概念框架来研究这个问题,并有助于设计系统来解决大规模的表发现和对齐问题。目前,针对海量表库的数据发现解决方案还处于起步阶段。有些解决方案高度依赖于特定领域。例如,用于在海量协作数据(通常称为网络表格)中查找相关表格的解决方案可以假设表格被设计用于具有丰富的、人类可读的属性名称或元数据的人类消费,并且相对较小(被设计为在网页上显示)。此外,解决方案通常假设数据科学家非常了解哪些数据是可用的,以及他们希望如何将其与已知数据进行集成。这些解决方案允许用户查找与指定属性联接的表或与查询表的联合。但是,如果扩展查询表的最佳方法是在几个属性上实际将其与其他两个表连接,然后将扩展结果与现有的更宽的表合并,则它们是不够的。这个项目将开发一种更全面的表发现方法,既可以发现一组可对齐的表,也可以找到将新数据与查询表集成(或对齐)的最佳方法。在这种称为“表即查询”的新范例中,用户不需要先验地知道存储库中的各种表在哪些属性上最适合。该项目促进了一项研究议程,在该议程下,Discovery查找的不是单个表,而是一组可以与查询表组合(对齐)的表。解决方案将包括表发现过程中的集成选择,查找与查询表最匹配的一组表,并找出最佳对齐方式。重要的是,该项目将不依赖于唯一名称假设,该假设规定不同的值指的是不同的唯一实体。真实数据包含同义词(两个引用同一实体的值)和同形异义词(一个值引用多个实体)。该项目将为研究表格对齐和发现确定新的基础和数学原理。搜索空间很大,因此该项目还将开发近似的、可扩展的解决方案,可以快速(以交互速度)在拥有数百万张表格的大型表格存储库中找到一组良好的表格和良好的比对。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Fueled by advances in information extraction and societal trends that value institutional openness and transparency, structured data are being produced and shared at an overwhelming speed. Open data sharing is central to supporting institutional transparency, but transparency is not achieved if shared data cannot be found and effectively aligned with other data being studied by data scientists, journalists, and others. This project will fundamentally contribute to the new science of open data sharing. The requirements for data discovery and integration over heterogeneous table repositories containing structured data are fundamentally different than they are for federated data integration (where for example, all data within an enterprise is integrated) or data exchange (where data is exchanged among a small set of autonomous peers, for example, between two institutions). This project will lay the theoretical foundations of data discovery (identification, alignment, and integration of tables) within table repositories. It will contribute both to developing the right conceptual framework for studying this problem and to designing systems that solve the table discovery and alignment problems at scale.Today, solutions for data discovery over massive table repositories are in their infancy. Some solutions are highly tied to a specific domain. For example, solutions for finding relevant tables in mass collaboration data (often called web tables) may assume tables are designed for human consumption with rich, human-readable attribute names or metadata, and are relatively small (being designed for display on web pages). Furthermore, solutions often assume that the data scientists know a lot about what data is available and exactly how they want to integrate it with known data. These solutions let a user find tables that join with a specified attribute or union with a query table. But they are inadequate if the best way to extend a query table is to actually join it on several attributes with two other tables and then union the extended result with an existing wider table. This project will develop a more holistic approach to table discovery that both discovers a set of alignable tables as well as the best way to integrate (or align) the new data with a query table. In this new paradigm called "table-as-query", the user does not need to know a priori on which attributes various tables in a repository are best aligned. This project promotes a research agenda under which discovery finds not a single table, but a set of tables that can be combined (aligned) with the query table. The solutions will include integration choices within the table discovery process, looking for a set of tables that can best be aligned with a query table and also finding what the best alignment is. Importantly, the project will not rely on the unique name assumption, which states that different values refer to different and unique entities. Real data contains synonyms (two values that refer to the same entity) and homographs (one value that refers to more than one entity). This project will define new foundations and mathematical principles for studying table alignment and discovery. The search space is massive, so the project will also develop approximate, scalable solutions that can quickly (at interactive speeds) find a good set of tables and good alignments over massive table repositories with millions of tables.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
Tractable Orders for Direct Access to Ranked Answers of Conjunctive Queries
用于直接访问连接查询的排名答案的易于处理的顺序
DOI: 10.1145/3578517
发表时间: 2023
期刊: ACM Transactions on Database Systems
影响因子: 1.8
作者: [Carmeli, Nofar, Tziavelis, Nikolaos, Gatterbauer, Wolfgang, Kimelfeld, Benny, Riedewald, Mirek]
通讯作者: Riedewald, Mirek
DOI: 10.5441/002/edbt.2021.03
发表时间: 2021-03
期刊: ArXiv
影响因子: --
作者: [Aristotelis Leventidis;Laura Di Rocco;Wolfgang Gatterbauer;Renée J. Miller;Mirek Riedewald]
通讯作者: Aristotelis Leventidis;Laura Di Rocco;Wolfgang Gatterbauer;Renée J. Miller;Mirek Riedewald
SANTOS: Relationship-based Semantic Table Union Search
SANTOS:基于关系的语义表联合搜索
DOI: 10.1145/3588689
发表时间: 2023
期刊: Proceedings of the ACM on Management of Data
影响因子: --
作者: [Khatiwada, Aamod, Fan, Grace, Shraga, Roee, Chen, Zixuan, Gatterbauer, Wolfgang, Miller, Renée J., Riedewald, Mirek]
通讯作者: Riedewald, Mirek
DIALITE: Discover, Align and Integrate Open Data Tables
DIALITE:发现、调整和集成开放数据表
DOI: 10.1145/3555041.3589732
发表时间: 2023
期刊: ACM SIGMOD
影响因子: --
作者: [Khatiwada, Aamod, Shraga, Roee, Miller, Renée J.]
通讯作者: Miller, Renée J.
共 10 条
    III: Small: Semantic Version Management in Data Lakes
    • 批准号:
      2325632
    • 项目类别:
      Standard Grant
    • 资助金额:
      $60.0万
    • 财政年份:
      2023
    • 负责人:
      Renee Miller
    • 依托单位:
    III : Medium: Collaborative Research: From Open Data to Open Data Curation
    • 批准号:
      2107248
    • 项目类别:
      Standard Grant
    • 资助金额:
      $48.0万
    • 财政年份:
      2021
    • 负责人:
      Renee Miller
    • 依托单位:
    CAREER: Managing Schematic Heterogeneity in Database Management Systems
    海外基金