课题基金 / 基金详情

EAGER: Collaborative Research: Mining Scientific Literature with the LAPPS Grid

EAGER: Collaborative Research: Mining Scientific Literature with the LAPPS Grid
EAGER:协作研究:使用 LAPPS 网格挖掘科学文献
批准号:
1811402
负责人:
James Pustejovsky
金额:
$9.93万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-01 至 2019-12-31

项目摘要

项目成果

James Pustejovsky的其他基金

相似基金

相关文献

中文摘要
翻译
科学家们已经无法跟上不断增长的科学出版物的数量。缺乏这种能力是科学进步的根本瓶颈。目前的搜索技术是有限的,因为它们能够找到许多相关的文档,但不能提取和组织这些文档的信息内容,也不能根据组织好的内容提出新的科学假设。基于自然语言处理(NLP)的文本挖掘策略是解决这一问题的公认方法,但大多数科学家没有专业知识或时间来使用它们。此外,NLP工具之间缺乏互操作性以及分散在网络上的存储库中的数据是共享工作流、资源和结果的障碍。该项目将确定在一个易于使用的挖掘科学文本的平台中需要哪些分析特性,实现这样一个平台的初始版本,并使其可供科学家使用。目前还没有一个开放的、易于使用的平台来挖掘科学文本,为广泛的软件、计算资源和出版数据提供可互操作的访问。公开可用的软件(如谷歌)并不面向发布数据,内部工具也很脆弱,只能提供一小部分相关结果。因此,该项目的主要目标是:(1)确定从科学出版物中挖掘信息的易于使用的平台的需求;(2)部署满足这些需求的设施。为了实现这一目标,该项目将扩展已经存在的nsf资助的LAPPS网格,包括访问广泛的可互操作的NLP工具、大量的出版数据、词汇和本体论资源的方法,并且,至关重要的是,将现有软件快速适应新的领域并评估结果。该项目还将利用NSF资助的Galaxy平台的增强功能,用于交互式数据探索和扩展对NSF硬件资源(XSEDE机器,包括Stampede、Bridges和Jetstream)的访问。通过提供获取科学出版物的服务,降低因许可、再分发和知识产权问题而产生的进入壁垒,该项目提供了科学家以前无法获得的能力。研究人员能够通过基于web的界面使用HPC基础设施执行大规模文本挖掘,而无需了解底层基础设施。此外,提供迭代的领域适应能力使科学家能够轻松地将现有服务适应专门领域,而无需配置或安装额外的组件。检查分散在大量出版物存储库中的显性和隐性信息的能力无疑将产生新的观察和见解。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Scientists have become unable to keep up with the ever-expanding number of scientific publications. The lack of this ability is a fundamental bottleneck to scientific progress. Current search technologies are limited because they are able to find many relevant documents, but cannot extract and organize the information content of these documents or suggest new scientific hypotheses based on the organized content. Natural Language Processing (NLP) based text mining strategies are a recognized means to approach this problem, but most scientists do not have the expertise or time to take use them. In addition, the lack of interoperability among NLP tools as well as the data in repositories scattered around the web are barriers to sharing workflows, resources, and results. This project will identify what analysis features are needed within an easy-to-use platform for mining scientific texts, implement an initial version of such a platform, and make it available to scientists.There is currently no open, easy-to-use platform for mining scientific texts that provides interoperable access to a wide array of software, computing resources, and publication data. Publicly available software (such as Google) is not geared toward publication data, and in-house tools are fragile and deliver only a fraction of relevant results. The main objective of this project is, therefore, to (1) identify the requirements for an easy-to-use platform for mining information from scientific publications and (2) deploy facilities that meet these needs. To achieve this goal this project will extend the already existing NSF-funded LAPPS Grid to include means to access a broad range of interoperable NLP tools, large bodies of publication data and lexical and ontological resources, and, crucially, to rapidly adapt existing software to new domains and evaluate results. This project will also leverage enhancements to the NSF-funded Galaxy platform for interactive data exploration and extended access to NSF hardware resources (XSEDE machines including Stampede, Bridges, and Jetstream). By providing access to services for mining scientific publications and lowering the barriers to entry resulting from licensing, redistribution, and intellectual property concerns, this project provides capabilities that were previously unavailable to scientists. Researchers are able to perform large-scale text mining using an HPC infrastructure through a web-based interface without the need to know about underlying infrastructure. Additionally, providing iterative domain adaptation capabilities enables scientists to easily adapt existing services to specialized areas without configuring or installing additional components. The ability to examine both explicit and implicit information scattered across massive repositories of publications will undoubtedly result in new observations and insights.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Integrating Dense Paraphrased-Enriched Representations with Large Language Models
  • 批准号:
    2326985
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2023
  • 负责人:
    James Pustejovsky
  • 依托单位:
Elements: Towards a Robust Cyberinfrastructure for NLP-based Search and Discoverability over Scientific Literature
  • 批准号:
    2104025
  • 项目类别:
    Standard Grant
  • 资助金额:
    $39.96万
  • 财政年份:
    2021
  • 负责人:
    James Pustejovsky
  • 依托单位:
Travel Support for North American Summer School for Logic, Language, and Information (NASSLLI)
  • 批准号:
    2002141
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.9万
  • 财政年份:
    2020
  • 负责人:
    James Pustejovsky
  • 依托单位:
Collaborative Research: NSF2026: EAGER: A Playground and Proposal for Growing an AGI
  • 批准号:
    2033932
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2020
  • 负责人:
    James Pustejovsky
  • 依托单位:
海外基金