课题基金 / 基金详情

NSF Convergence Accelerator Track D: The Data Hypervisor: Orchestrating Data and Models

NSF Convergence Accelerator Track D: The Data Hypervisor: Orchestrating Data and Models
NSF 融合加速器轨道 D:数据管理程序:编排数据和模型
批准号:
2040718
负责人:
Ian Foster
金额:
$95.46万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-15 至 2023-05-31

项目摘要

项目成果

Ian Foster的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
The NSF Convergence Accelerator supports use-inspired, team-based, multidisciplinary efforts that address challenges of national importance and will produce deliverables of value to society in the near future. This project, NSF Convergence Accelerator–Track D: The Data Hypervisor: Orchestrating Data and Models, will design and implement the Data Station—a new architecture where both data and derived data products are sealed and cannot be directly seen or downloaded by anyone. In the Data Station architecture, computation is brought to the data, rather than data being brought to users, as is common in traditional data lakes and warehouses. Sharing data and models has had a transformative impact on scientific problems from medical imaging to natural language understanding. Despite the potential upside, many researchers in both academia and industry are reluctant to centralize and share data to both internal and external researchers. Organizations today have to navigate complex regulatory considerations and protect intellectual property while incurring a significant technical investment in documenting and maintaining data. The Data Station will ease access to sensitive data, assist with data discovery and integration, and facilitate enforcement of arbitrary data access and governance policies. The project will work with partners in biomedicine, materials science, and enterprise data management to establish the capabilities and prove the concepts of the Data Station architecture.While building upon prior research in data systems, Data Station will introduce novel data-unaware task capsules that enable users to specify data-driven tasks such as traditional data queries and machine learning model training without the user requiring direct access to the data itself. The programming interfaces convey sufficient information for the Data Station to trigger the discovery of potentially relevant datasets; integrate and prune those datasets for computation; and compute the results by executing the task. In effect, Data Station inverts the traditional data querying modeling by bringing computations to the data. Task capsules also include a user-defined metric for determining what results are useful from the user’s perspective as well as which trust constraints need to be met to validate the provenance of input datasets. The Data Station captures metadata every time a derived data product is created and provides a set of primitives to implement various data governance and data access policies necessary to address data contributor use cases. Only authorized users are able to access the data based on a novel access-token model implemented by Data Station that permits fine-grained yet scalable access control. Users must explicitly be authorized to access results via tokens obtained from data contributors. The Data Station project will engage a diverse set of partners in materials science, biomedicine, and enterprise scenarios to help design and apply the Data Station to various use cases. An education program will engage high school, undergraduate, and graduate students in researching, developing, and evaluating the Data Station.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1002/aaai.12042
发表时间: 2022-03-01
期刊: AI MAGAZINE
影响因子: 0.9
作者: [Baru, Chaitanya, Pozmantier, Michael, Zhang, Peng]
通讯作者: Zhang, Peng
DOI: 10.1109/icde55515.2023.00045
发表时间: 2021-06
期刊: 2023 IEEE 39th International Conference on Data Engineering (ICDE)
影响因子: --
作者: [Yue Gong;Zhiru Zhu;Sainyam Galhotra;R. Fernandez]
通讯作者: Yue Gong;Zhiru Zhu;Sainyam Galhotra;R. Fernandez
Data-Sharing Markets: Model, Protocol, and Algorithms to Incentivize the Formation of Data-Sharing Consortia
数据共享市场:激励数据共享联盟形成的模型、协议和算法
DOI: --
发表时间: 2023
期刊: Proceedings ACMSIGMOD International Conference on Management of Data
影响因子: --
作者: [Raul Castro Fernandez]
通讯作者: Raul Castro Fernandez
DOI: 10.14778/3551793.3551861
发表时间: 2022-07
期刊: ArXiv
影响因子: --
作者: [Siyuan Xia;Zhiru Zhu;Chris Zhu;Jinjin Zhao;K. Chard;Aaron J. Elmore;Ian D. Foster;Michael]
通讯作者: Siyuan Xia;Zhiru Zhu;Chris Zhu;Jinjin Zhao;K. Chard;Aaron J. Elmore;Ian D. Foster;Michael
Collaborative Research: NSF Workshop on Automated, Programmable and Self Driving Labs
  • 批准号:
    2335910
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.2万
  • 财政年份:
    2023
  • 负责人:
    Ian Foster
  • 依托单位:
Frameworks: Garden: A FAIR Framework for Publishing and Applying AI Models for Translational Research in Science, Engineering, Education, and Industry
  • 批准号:
    2209892
  • 项目类别:
    Standard Grant
  • 资助金额:
    $349.65万
  • 财政年份:
    2022
  • 负责人:
    Ian Foster
  • 依托单位:
Collaborative Research: OAC Core: ScaDL: New Approaches to Scaling Deep Learning for Science Applications on Supercomputers
  • 批准号:
    2107511
  • 项目类别:
    Standard Grant
  • 资助金额:
    $27.16万
  • 财政年份:
    2021
  • 负责人:
    Ian Foster
  • 依托单位:
Collaborative Research: Frameworks: funcX: A Function Execution Service for Portability and Performance
  • 批准号:
    2004894
  • 项目类别:
    Standard Grant
  • 资助金额:
    $265.81万
  • 财政年份:
    2020
  • 负责人:
    Ian Foster
  • 依托单位:
海外基金