课题基金 / 基金详情

III: Medium: Table-as-Query: Unifying Data Discovery and Alignment

III: Medium: Table-as-Query: Unifying Data Discovery and Alignment
III:媒介:表即查询:统一数据发现和对齐
批准号:
1956096
负责人:
Renee Miller
金额:
$100.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-08-01 至 2024-07-31

项目摘要

项目成果

Renee Miller的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Fueled by advances in information extraction and societal trends that value institutional openness and transparency, structured data are being produced and shared at an overwhelming speed. Open data sharing is central to supporting institutional transparency, but transparency is not achieved if shared data cannot be found and effectively aligned with other data being studied by data scientists, journalists, and others. This project will fundamentally contribute to the new science of open data sharing. The requirements for data discovery and integration over heterogeneous table repositories containing structured data are fundamentally different than they are for federated data integration (where for example, all data within an enterprise is integrated) or data exchange (where data is exchanged among a small set of autonomous peers, for example, between two institutions). This project will lay the theoretical foundations of data discovery (identification, alignment, and integration of tables) within table repositories. It will contribute both to developing the right conceptual framework for studying this problem and to designing systems that solve the table discovery and alignment problems at scale.Today, solutions for data discovery over massive table repositories are in their infancy. Some solutions are highly tied to a specific domain. For example, solutions for finding relevant tables in mass collaboration data (often called web tables) may assume tables are designed for human consumption with rich, human-readable attribute names or metadata, and are relatively small (being designed for display on web pages). Furthermore, solutions often assume that the data scientists know a lot about what data is available and exactly how they want to integrate it with known data. These solutions let a user find tables that join with a specified attribute or union with a query table. But they are inadequate if the best way to extend a query table is to actually join it on several attributes with two other tables and then union the extended result with an existing wider table. This project will develop a more holistic approach to table discovery that both discovers a set of alignable tables as well as the best way to integrate (or align) the new data with a query table. In this new paradigm called "table-as-query", the user does not need to know a priori on which attributes various tables in a repository are best aligned. This project promotes a research agenda under which discovery finds not a single table, but a set of tables that can be combined (aligned) with the query table. The solutions will include integration choices within the table discovery process, looking for a set of tables that can best be aligned with a query table and also finding what the best alignment is. Importantly, the project will not rely on the unique name assumption, which states that different values refer to different and unique entities. Real data contains synonyms (two values that refer to the same entity) and homographs (one value that refers to more than one entity). This project will define new foundations and mathematical principles for studying table alignment and discovery. The search space is massive, so the project will also develop approximate, scalable solutions that can quickly (at interactive speeds) find a good set of tables and good alignments over massive table repositories with millions of tables.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
Tractable Orders for Direct Access to Ranked Answers of Conjunctive Queries
用于直接访问连接查询的排名答案的易于处理的顺序
DOI: 10.1145/3578517
发表时间: 2023
期刊: ACM Transactions on Database Systems
影响因子: 1.8
作者: [Carmeli, Nofar, Tziavelis, Nikolaos, Gatterbauer, Wolfgang, Kimelfeld, Benny, Riedewald, Mirek]
通讯作者: Riedewald, Mirek
DOI: 10.5441/002/edbt.2021.03
发表时间: 2021-03
期刊: ArXiv
影响因子: --
作者: [Aristotelis Leventidis;Laura Di Rocco;Wolfgang Gatterbauer;Renée J. Miller;Mirek Riedewald]
通讯作者: Aristotelis Leventidis;Laura Di Rocco;Wolfgang Gatterbauer;Renée J. Miller;Mirek Riedewald
SANTOS: Relationship-based Semantic Table Union Search
SANTOS:基于关系的语义表联合搜索
DOI: 10.1145/3588689
发表时间: 2023
期刊: Proceedings of the ACM on Management of Data
影响因子: --
作者: [Khatiwada, Aamod, Fan, Grace, Shraga, Roee, Chen, Zixuan, Gatterbauer, Wolfgang, Miller, Renée J., Riedewald, Mirek]
通讯作者: Riedewald, Mirek
DIALITE: Discover, Align and Integrate Open Data Tables
DIALITE:发现、调整和集成开放数据表
DOI: 10.1145/3555041.3589732
发表时间: 2023
期刊: ACM SIGMOD
影响因子: --
作者: [Khatiwada, Aamod, Shraga, Roee, Miller, Renée J.]
通讯作者: Miller, Renée J.
10
    III: Small: Semantic Version Management in Data Lakes
    • 批准号:
      2325632
    • 项目类别:
      Standard Grant
    • 资助金额:
      $60.0万
    • 财政年份:
      2023
    • 负责人:
      Renee Miller
    • 依托单位:
    III : Medium: Collaborative Research: From Open Data to Open Data Curation
    • 批准号:
      2107248
    • 项目类别:
      Standard Grant
    • 资助金额:
      $48.0万
    • 财政年份:
      2021
    • 负责人:
      Renee Miller
    • 依托单位:
    CAREER: Managing Schematic Heterogeneity in Database Management Systems
    海外基金