课题基金 / 基金详情

III: Medium: Collaborative Research: Extracting and Linking AI Artifacts

III: Medium: Collaborative Research: Extracting and Linking AI Artifacts
III:媒介:协作研究:提取和链接人工智能工件
批准号:
2107213
负责人:
Eduard Dragut
金额:
$67.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-12-01 至 2024-11-30

项目摘要

项目成果

Eduard Dragut的其他基金

相似基金

相关文献

中文摘要
翻译
该项目的目标是创建一个框架,用于连接人工智能(AI)工作流的所有突出方面,包括数据、AI模型、AI工具、任务和培训方法。研究人员试图创建一个框架,从整体上看待人工智能工作流程,从而为科学办公室人工智能数据圆桌会议报告中确定的三个关键问题之一提供解决方案:“用相关数据、模型和任务的框架解决人工智能中的开放性问题。”联邦资助机构的关键条款之一是创建和公开传播研究工件(例如,数据、模型)。尽管基于出版物的知识很容易重用,但数据和模型却不能。数据是生成人工智能模型的关键要素。然而,人工智能模型和用于生成它或它解决的任务的数据之间的关系,以及人工智能模型所测试的数据之间的关系,既不是由模型捕获的,也不是由数据或任务捕获的。因此,研究者试图建立一种统一的方法来构建这种关系并对其进行注释。这个项目将有助于广泛的信息检索领域,特别是已命名实体识别领域。在这个项目中,命名的实体是数据集、AI模型、开发工具和各种方法的名称,例如在培训中使用的方法。研究人员将采用一种整体方法来管理人工智能研究工件,即纸张-任务-数据-模型-工具,这反过来将产生一种创新的方式来概念化和执行数据-人工智能模型搜索和聚合。该项目的技术创新在于创造了实体和关系提取以及实体链接的新技术。该项目还将为科学文献挖掘领域做出贡献。研究人员将创造新的技术来自动识别和编目公共人工智能数据和模型,以提高其可重用性。关键的洞察力是,没有研究论文本身,研究人工智能工件缺乏必要的重用上下文。例如,论文描述了数据集的作用(例如,训练或测试),并告诉模型是原始的还是用作基线的。通过自动推断任务-数据-模型关系,该项目将增加向新企业推荐工件的能力,从而缩短相关工件搜索的时间。在教育方面,这项工作将包括培训研究生和本科生,特别是鼓励妇女和代表性不足的群体参与研究工作和课程编制。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The goal of this project is to create a framework for linking all salient aspects of an artificial intelligence (AI) workflow, including data, AI models, AI tools, tasks, and training methodology. The investigators seek to create a framework that takes a holistic view of the AI workflow, and thus, will provide a solution to one of the three key problems identified in the Report of the Office of Science Roundtable on Data for AI: “Address open questions in AI with frameworks for relating data, models, and tasks.” One of the key provisions of federal funding agencies is the creation and open dissemination of research artifacts (e.g., data, models). Although publication-based knowledge is easily reused, data and models are not. Data are the key ingredients to generate AI models. However, the relation between an AI model and the data used to generate it or the task it solves, and the data on which the AI model is tested on, is captured by neither the model nor the data or task. Thus, the investigators seek to create a unified approach to construct this relationship and annotate it. This project will contribute to the broad field of information retrieval and, in particular, to the field of named entity recognition. In this project, the named entities are the datasets, AI models, developing tools, and the names of various methods, such as those employed in training. The investigators will employ a holistic approach to the management of AI research artifacts, i.e., paper-task-data-model-tool, which in turn will produce an innovative way to conceptualize and execute data-AI model search and aggregation. The technical innovation of this project is the creation of novel techniques for entity and relation extraction as well as for entity linking. The project will also contribute to the field of scientific literature mining. The investigators will create novel technology to automatically identify and catalog public AI data and models that increase their reusability. The key insight is that, without the research papers themselves, the research AI artifacts lack the necessary context for reuse. For example, papers describe the role of a dataset (e.g., training or testing) and tell if a model is original or used as a baseline. By automatically inferring task-data-model relations, this project will increase the ability of suggesting artifacts to a new undertaking, thus shortening the time for relevant artifact search. Educationally, this work will involve training of graduate and undergraduate students, particularly encouraging the participation of women and underrepresented groups in the research efforts, and curriculum development.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1162/tacl_a_00592
发表时间: 2023-05
期刊: Transactions of the Association for Computational Linguistics
影响因子: 10.9
作者: [Huitong Pan;Qi Zhang;E. Dragut;Cornelia Caragea;Longin Jan Latecki]
通讯作者: Huitong Pan;Qi Zhang;E. Dragut;Cornelia Caragea;Longin Jan Latecki
Proto-OKN Theme 1: Knowledge Graph to Support Evaluation and Development of Climate Models
  • 批准号:
    2333789
  • 项目类别:
    Cooperative Agreement
  • 资助金额:
    $149.86万
  • 财政年份:
    2023
  • 负责人:
    Eduard Dragut
  • 依托单位:
NSF Convergence Accelerator Track F: America's Fourth Estate at Risk: A System for Mapping the (Local) Journalism Life Cycle to Rebuild the Nation's News Trust
  • 批准号:
    2137846
  • 项目类别:
    Standard Grant
  • 资助金额:
    $75.0万
  • 财政年份:
    2021
  • 负责人:
    Eduard Dragut
  • 依托单位:
BIGDATA: F: Collaborative Research: Collective Mining of Vertical Social Communities
  • 批准号:
    1838145
  • 项目类别:
    Standard Grant
  • 资助金额:
    $42.79万
  • 财政年份:
    2018
  • 负责人:
    Eduard Dragut
  • 依托单位:
BIGDATA: Collaborative Research: F: Streaming Architecture for Continuous Entity Linking in Social Media
  • 批准号:
    1546480
  • 项目类别:
    Standard Grant
  • 资助金额:
    $78.33万
  • 财政年份:
    2016
  • 负责人:
    Eduard Dragut
  • 依托单位:
海外基金