课题基金 / 基金详情

Collaborative Research: IIS: III: MEDIUM: Learning Protein-ish: Foundational Insight on Protein Language Models for Better Understanding, Democratized Access, and Discovery

Collaborative Research: IIS: III: MEDIUM: Learning Protein-ish: Foundational Insight on Protein Language Models for Better Understanding, Democratized Access, and Discovery
协作研究:IIS:III:中等:学习蛋白质:对蛋白质语言模型的基础洞察,以更好地理解、民主化访问和发现
批准号:
2310113
负责人:
Amarda Shehu
金额:
$59.99万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-01 至 2026-07-31

项目摘要

项目成果

Amarda Shehu的其他基金

相似基金

相关文献

中文摘要
翻译
大型语言模型是大规模的神经网络,它学习单词的丰富上下文表示,并使用这些表示来处理自然语言处理(NLP)中的各种任务。这些模型是生成式人工智能的一个突出例子,正在成为提取和组织海量生物数据库的内容以及预测广泛的分子生物特性的有前途的方法。然而,令人惊讶的是,我们对这些模型在其学习表示中捕捉到了什么,为什么它们在一些任务上表现良好而在另一些任务中表现不佳,以及它们如何能够对描述生物空间的关系产生深刻的洞察,我们知之甚少。如果NLP方面的进展是一个迹象,那么目前通过大幅增加语言模型的可训练参数数量来改善其性能的趋势是不可持续的,无论是对于我们的碳足迹,还是对于确保学术环境中研究和学术的公平/可获得性而言。该项目使用原理蛋白质语言模型(PLM)作为计算工具,在信息集成和信息学的交叉点推进算法研究,以更深入地了解不同细节和规模的蛋白质空间中的结构、功能和进化组织。它还旨在以一种资源意识强、可持续和所有研究人员都能接触到的方式来实现这一目标。研究活动被组织在三个方面:(1)在PLM中编码先前的生物知识,以便在复合空间中进行联合和资源感知学习;(2)揭示基本属性并组织学习的表示空间,以告知并连接所捕获的内容和感兴趣的属性;以及(3)使PLM能够捕获不同的上下文,以便更深入地探索蛋白质空间中的结构、功能和进化组织。这种跨学科的方法有助于机器学习、生物信息学和分子生物学领域,并在这些学科的交界处为所有级别的代表不足的学生提供培训机会。研究人员决心在社区和学科之间架起桥梁,他们计划了一些活动,以建立和激励一个跨学科的社区,以进一步推进他们的研究。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Large language models are massive neural networks that learn rich contextual representations of words and use such representations to address a variety of tasks in natural language processing (NLP). These models are a prominent example of generative artificial intelligence and are emerging as promising approaches for distilling and organizing the content of massive biological databases and for predicting a wide range of molecular bio-properties. Yet, we know surprisingly little about what these models capture in their learned representations, why they perform well on some tasks and not on others, and how they can produce deep insight into the relationships describing the biological space. If progress in NLP is any indication, the current trend of improving the performance of language models by drastically increasing the number of their trainable parameters is unsustainable both for our carbon footprint and for ensuring equity/accessibility of research and scholarship in the academic setting. This project advances algorithmic research at the intersection of information integration and informatics using principled protein language models (PLMs) as computational vehicles for deeper insight into the structural, functional, and evolutionary organization across protein space at varying levels of detail and scale. It also aims to do so in a way that is resource-aware, sustainable, and accessible to all researchers. The research activities are organized in three thrusts: (1) encoding prior biological knowledge in PLMs for joint and resource-aware learning in composite spaces, (2) revealing fundamental properties and organizing the learned representation space to inform and connect what is captured with properties of interest, and (3) enabling PLMs to capture diverse contexts for deeper exploration of the structural, functional, and evolutionary organization across protein space. This interdisciplinary approach contributes to the fields of machine learning, bioinformatics, and molecular biology and provides opportunities at the interface of these disciplines for training under-represented students of all levels. The investigators are determined to bridge communities and disciplines, and they have planned activities to build and galvanize a trans-disciplinary community to further advance their research.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Conference: Large Language Models for Biological Discoveries (LLMs4Bio)
  • 批准号:
    2411529
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.95万
  • 财政年份:
    2024
  • 负责人:
    Amarda Shehu
  • 依托单位:
Collaborative Research: IIBR: Innovation: Bioinformatics: Linking Chemical and Biological Space: Deep Learning and Experimentation for Property-Controlled Molecule Generation
  • 批准号:
    2318829
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $29.93万
  • 财政年份:
    2023
  • 负责人:
    Amarda Shehu
  • 依托单位:
Intergovernmental Personnel Act
  • 批准号:
    1948645
  • 项目类别:
    Intergovernmental Personnel Award
  • 资助金额:
    $21.51万
  • 财政年份:
    2019
  • 负责人:
    Amarda Shehu
  • 依托单位:
Collaborative: SI2-SSE - A Plug-and-Play Software Platform of Robotics-Inspired Algorithms for Modeling Biomolecular Structures and Motions
  • 批准号:
    1440581
  • 项目类别:
    Standard Grant
  • 资助金额:
    $21.73万
  • 财政年份:
    2015
  • 负责人:
    Amarda Shehu
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)