III: Small: Semantic Version Management in Data Lakes
III: Small: Semantic Version Management in Data Lakes
批准号:
2325632
负责人:
Renee Miller
金额:
$60.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-15 至 2026-08-31
中文摘要
数据为我们的经济提供动力。那些与数据打交道的人将他们的努力投入到寻找有效的方法来提取知识和处理其日益增长的规模上。这种增长不仅表现为新的数据源,而且还表现为数据复制-即复制、集成和修改数据集,从而创建新版本的数据集。值得注意的是,国际数据公司(International Data Corporation)等咨询公司估计,商业中使用的大多数新生成的数据都是现有数据的版本。因此,理解数据需要对数据版本化的语义理解,这成为处理和管理数据的关键因素。该项目将专注于促进对数据版本控制的科学理解,并将从根本上促进任何使用数据的科学或活动,目前数据涵盖了所有人类活动的大量内容。为了推进开放数据科学,该项目将为理解导致数据新版本的数据变化的语义奠定基础,并将引入可扩展的工具来发现和解释数据变化。这将有助于制定有效的框架,在数据科学管道内处理多种数据版本,并有助于设计纳入和管理数据复制的系统。这项工作预计还将通过促进负责任和开放的数据科学而造福社会。它的解决方案将公之于众,并与高质量、高度精心策划的基准一起提供,这些基准本身将在进行比较和解决科学辩论方面具有科学价值,以推动这一重要领域的发展。该项目还将使用“负责任的数据科学”的各个方面,旨在确保在处理数据时公平、准确、保密和透明。这个项目将开发一种新的范例,我们称之为语义版本管理。我们的愿景是使用户能够以最小的前期工作了解通常驻留在数据湖中的众多版本。该项目的主要目标是使目前主要依赖文件名的数据科学家能够找到数据集的“正确”版本,以查看和理解数据集之间所做的更改(清理、价值分配、集成等)。该研究方法建立、集成和扩展了可伸缩数据发现、逐例编程和数据转换合成以及从不一致和不完整的证据中学习模式映射的工作。该项目将开发支持数据版本化语义理解的方法,为研究数据版本奠定基础,并建立评估和基准数据版本化的新方法。具体地说,这个项目将解决以下基本研究挑战:1)恢复对数据进行的转换,并解释一个数据集与另一个数据集版本的不同之处;2)从海量表库或数据湖中高效地找到数据集的版本;以及3)了解版本集合中的版本历史,并构建一个图表来表达数据版本创建背后的故事。在整个开发过程中,该项目还将开发新的评估框架,不仅考虑解决方案的正确性,还考虑其可解释性。语义版本管理的一个重要动机是让用户对他们使用的数据有更多的信任。如果他们了解用于从一个版本派生另一个版本的转换,他们就可以更好地了解一个版本是否满足他们的数据科学任务的需求。此外,将生成新的基准并与社区共享,以鼓励开放科学,并允许与新的或替代的版本理解方法进行可靠的比较。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Data fuels our economy. Those who work with data invest their efforts in finding effective ways to extract knowledge and to handle its increasing size. This growth is not only characterized by new sources of data, but also by data replication - that is, the copying, integration, and modification of datasets that creates new versions of datasets. Notably, consultants such as the International Data Corporation estimate that most of the newly generated data being used in business are versions of existing data. Understanding data therefore requires a semantic understanding of data versioning, which becomes a key ingredient in handling and managing data. This project will focus on advancing the scientific understanding of data versioning and will fundamentally contribute to any science or activity that uses data, which nowadays covers a tremendous amount of all human activity. To advance open data science, this project will lay the foundations for semantic understanding of data changes that result in new versions of data and will introduce scalable tools to uncover and explain data changes. This will contribute both to the development of effective frameworks to handle multiple data versions within a data science pipeline as well as to the design of systems that incorporate and manage data replication. This work is expected to also benefit society by facilitating responsible and open data science. Its solutions will be made publicly available and provided alongside high-quality highly curated benchmarks that themselves will have scientific value in allowing comparisons and settling scientific debates in order to advance this important field. The project will also use aspects of "responsible data science", aiming to ensure fairness, accuracy, confidentiality, and transparency when working with data. This project will develop a new paradigm we call semantic version management. The vision is to enable users, with minimal upfront effort, to understand the multitude of versions that typically reside in data lakes. The main objective of the project is to enable data scientists who currently rely mainly on file names to find the "right" version of a dataset to see and understand the changes (cleaning, value imputation, integration, and others) that have been made between datasets. The research methodology builds-on, integrates, and extends work on scalable data discovery; program by example and data transformation synthesis; and learning schema mappings from inconsistent and incomplete evidence. This project will develop methods to support the semantic understanding of data versioning, lay the foundations for studying data versions, and establish new methods for evaluating and benchmarking data versioning. Specifically, this project will address the following fundamental research challenges: 1) recovering transformations done to data and explain how one dataset differs from another version of the dataset; 2) efficiently finding versions of a dataset from within a massive table repository or data lake; and 3) understanding the version history among a collection of versions and constructing a graph that expresses the story behind the creation of the data versions. Throughout the development, this project will also develop new evaluation frameworks that not only consider the correctness of solutions, but also their explainability. An important motivation for semantic version management is to give users more trust in the data they are using. If they understand the transformations used to derive one version from another, they can better understand if a version meets the needs of their data science task. In addition, new benchmarks will be generated and shared with the community to encourage open science and allow reliable comparison with new or alternative approaches to version understanding.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III : Medium: Collaborative Research: From Open Data to Open Data Curation
-
批准号:2107248
-
项目类别:Standard Grant
-
资助金额:$48.0万
-
财政年份:2021
-
负责人:Renee Miller
-
依托单位:
III: Medium: Table-as-Query: Unifying Data Discovery and Alignment
-
批准号:1956096
-
项目类别:Continuing Grant
-
资助金额:$100.0万
-
财政年份:2020
-
负责人:Renee Miller
-
依托单位:
CAREER: Managing Schematic Heterogeneity in Database Management Systems
-
批准号:9702974
-
项目类别:Continuing Grant
-
资助金额:$29.44万
-
财政年份:1997
-
负责人:Renee Miller
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: