NSF Convergence Accelerator Track D: Towards Intelligent Sharing and Search for AI Models and Datasets
NSF Convergence Accelerator Track D: Towards Intelligent Sharing and Search for AI Models and Datasets
批准号:
2040727
负责人:
Jingbo Shang
金额:
$94.72万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-15 至 2022-12-31
中文摘要
NSF融合加速器支持以使用为灵感、以团队为基础的多学科努力,以应对国家重要性的挑战,并将在不久的将来产生对社会有价值的成果。人工智能驱动的应用程序的一个主要目标是发现特定领域数据集中的潜在模式,这通常需要大量的现场经验和跨学科知识来设计甚至选择合适的人工智能模型。该项目将为人工智能数据集和模型开发一个枢纽和门户。它将提供数据和模型匹配建议,使用领域知识来改进数据集和模型的搜索策略,并支持隐私。中心和门户将吸引广泛的用户(在STEM和非STEM领域),在我们今天只能想象的各个领域创造人工智能驱动的创新。成功的执行将提供新的有形构件,包括模型和数据模式、软件、系统和服务,使人工智能模型和数据集易于发现、可访问、可互操作和可重现。将使用四种新技术来实现设想的系统:(1)具有自适应描述性统计的细粒度隐私控制技术,在数据所有者的隐私需求和应用程序驱动的可用性之间实现平衡。所有其他组件将只能访问受隐私控制的数据;(2)利用有关AI模型和数据集的各种信息(例如数据值、模型参数、辅助描述)将领域逻辑合并到语义中的自动元数据生成方法。将该元数据与模型和数据集一起组织为富文本网络;(3)将富文本网络中的信息转换到潜在空间的表示学习方法,在该潜在空间中,具有相似语义的数据集/模型将彼此接近。这种对多模式数据的学习将使人们能够全面理解模型和数据集;(4)将构建一个带约束的学习匹配模型,以连接数据集和模型。这些限制主要是由模型和数据集之间的模式对齐引起的,这也可以过滤掉明显不兼容的模型和数据集选择,显著加快搜索和匹配过程。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The NSF Convergence Accelerator supports use-inspired, team-based, multidisciplinary efforts that address challenges of national importance and will produce deliverables of value to society in the near future. A major goal of AI-driven applications is to discover the underlying patterns in domain-specific datasets, which typically requires tremendous field experience and interdisciplinary knowledge to design or even select suitable AI models. This project will develop a hub and portal for AI data sets and models. It will offer data and model matching recommendations, the use of domain knowledge to improve search strategies for data sets and models, and support for privacy. The hub and portal will engage a broad range of users (in STEM and non-STEM fields) creating AI-driven innovations in various domains that we can only imagine today. Successful execution will provide new tangible artifacts consisting of model and data schemas, software, systems, and services that would make the AI models and datasets easily discoverable, accessible, interoperable, and reproducible.Four novel techniques will be used to realize the envisioned system: (1) A fine-grained privacy control technique with adaptive descriptive statistics, achieving a balance between the privacy needs of data owners and application-driven usability. All other components will have access to only the privacy-controlled data; (2) An automated metadata generation method that exploits various kinds of information about AI models and datasets (e.g., data values, model parameters, auxiliary descriptions) to incorporate domain logic into semantics. This metadata, together with the models and datasets, will be organized as a text-rich network; (3) A representation learning method that transforms information in the text-rich network into a latent space, where datasets/models with similar semantics would be close to each other. This learning over multimodal data will enable comprehensive understandings about models and datasets; (4) A learning-to-match model with constraints will be built to bridge datasets and models. The constraints are mainly induced from schema alignment between models and datasets, which can also filter out obvious non-compatible model and dataset choices, significantly expediting the search and matching process.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(28)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
“Misc”-Aware Weakly Supervised Aspect Classification
–Misc – 感知弱监督方面分类
DOI:
--
发表时间:
2021
期刊:
Proceedings of the 2021 SIAM International Conference on Data Mining (SDM
影响因子:
--
作者:
[Li, Peiran, Guo, Fang, Shang, Jingbo]
通讯作者:
Shang, Jingbo
DOI:
10.1145/3534678.3539329
发表时间:
2022-08
期刊:
Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
作者:
[Ranak Roy Chowdhury;Xiyuan Zhang;Jingbo Shang;Rajesh K. Gupta;Dezhi Hong]
通讯作者:
Ranak Roy Chowdhury;Xiyuan Zhang;Jingbo Shang;Rajesh K. Gupta;Dezhi Hong
DOI:
10.48550/arxiv.2301.11459
发表时间:
2023-01
期刊:
ArXiv
影响因子:
--
作者:
[Zi Lin;J. Liu;Jingbo Shang]
通讯作者:
Zi Lin;J. Liu;Jingbo Shang
Sensei: Self-Supervised Sensor Name Segmentation
Sensei:自监督传感器名称分割
DOI:
10.18653/v1/2021.findings-acl.87
发表时间:
2021
期刊:
Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021
影响因子:
--
作者:
[Wu, Jiaman, Hong, Dezhi, Gupta, Rajesh, Shang, Jingbo]
通讯作者:
Shang, Jingbo
Sharing personal ECG time-series data privately
私下共享个人心电图时间序列数据
DOI:
10.1093/jamia/ocac047
发表时间:
2022
期刊:
Journal of the American Medical Informatics Association
影响因子:
6.4
作者:
[Bonomi, Luca, Wu, Zeyun, Fan, Liyue]
通讯作者:
Fan, Liyue
共 25 条
CAREER: Knowledge Extraction and Discovery from Massive Text Corpora via Extremely Weak Supervision
-
批准号:2239440
-
项目类别:Continuing Grant
-
资助金额:$60.0万
-
财政年份:2023
-
负责人:Jingbo Shang
-
依托单位:
海外基金