Collaborative Research: IIS: III: MEDIUM: Learning Protein-ish: Foundational Insight on Protein Language Models for Better Understanding, Democratized Access, and Discovery
Collaborative Research: IIS: III: MEDIUM: Learning Protein-ish: Foundational Insight on Protein Language Models for Better Understanding, Democratized Access, and Discovery
批准号:
2310113
负责人:
Amarda Shehu
金额:
$59.99万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-01 至 2026-07-31
中文摘要
大型语言模型是大规模的神经网络,它学习丰富的单词上下文表示,并使用这些表示来解决自然语言处理(NLP)中的各种任务。这些模型是生成式人工智能的一个突出例子,并且正在成为提取和组织大量生物数据库内容以及预测广泛的分子生物特性的有前途的方法。然而,我们对这些模型在其学习表征中捕获的内容知之甚少,为什么它们在某些任务上表现良好而在其他任务上表现不佳,以及它们如何对描述生物空间的关系产生深刻的见解。如果NLP的进展是任何迹象,那么当前通过大幅增加可训练参数数量来提高语言模型性能的趋势对于我们的碳足迹和确保学术环境中研究和奖学金的公平性/可及性来说都是不可持续的。该项目推进了信息集成和信息学交叉领域的算法研究,使用有原则的蛋白质语言模型(PLMs)作为计算工具,在不同的细节和规模水平上深入了解蛋白质空间的结构、功能和进化组织。它还旨在以一种资源意识、可持续和所有研究人员都可以使用的方式进行研究。研究活动分为三个重点:(1)在plm中编码先前的生物学知识,以便在复合空间中进行联合和资源感知学习;(2)揭示基本属性并组织学习到的表示空间,以便将捕获的内容与感兴趣的属性联系起来;(3)使plm能够捕获不同的背景,以便更深入地探索蛋白质空间中的结构、功能和进化组织。这种跨学科的方法为机器学习、生物信息学和分子生物学领域做出了贡献,并在这些学科的界面上为培养各级代表性不足的学生提供了机会。研究人员决心在社区和学科之间架起桥梁,他们计划开展活动,建立和激励一个跨学科的社区,以进一步推进他们的研究。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Large language models are massive neural networks that learn rich contextual representations of words and use such representations to address a variety of tasks in natural language processing (NLP). These models are a prominent example of generative artificial intelligence and are emerging as promising approaches for distilling and organizing the content of massive biological databases and for predicting a wide range of molecular bio-properties. Yet, we know surprisingly little about what these models capture in their learned representations, why they perform well on some tasks and not on others, and how they can produce deep insight into the relationships describing the biological space. If progress in NLP is any indication, the current trend of improving the performance of language models by drastically increasing the number of their trainable parameters is unsustainable both for our carbon footprint and for ensuring equity/accessibility of research and scholarship in the academic setting. This project advances algorithmic research at the intersection of information integration and informatics using principled protein language models (PLMs) as computational vehicles for deeper insight into the structural, functional, and evolutionary organization across protein space at varying levels of detail and scale. It also aims to do so in a way that is resource-aware, sustainable, and accessible to all researchers. The research activities are organized in three thrusts: (1) encoding prior biological knowledge in PLMs for joint and resource-aware learning in composite spaces, (2) revealing fundamental properties and organizing the learned representation space to inform and connect what is captured with properties of interest, and (3) enabling PLMs to capture diverse contexts for deeper exploration of the structural, functional, and evolutionary organization across protein space. This interdisciplinary approach contributes to the fields of machine learning, bioinformatics, and molecular biology and provides opportunities at the interface of these disciplines for training under-represented students of all levels. The investigators are determined to bridge communities and disciplines, and they have planned activities to build and galvanize a trans-disciplinary community to further advance their research.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Conference: Large Language Models for Biological Discoveries (LLMs4Bio)
-
批准号:2411529
-
项目类别:Standard Grant
-
资助金额:$1.95万
-
财政年份:2024
-
负责人:Amarda Shehu
-
依托单位:
Collaborative Research: IIBR: Innovation: Bioinformatics: Linking Chemical and Biological Space: Deep Learning and Experimentation for Property-Controlled Molecule Generation
-
批准号:2318829
-
项目类别:Continuing Grant
-
资助金额:$29.93万
-
财政年份:2023
-
负责人:Amarda Shehu
-
依托单位:
Intergovernmental Personnel Act
-
批准号:1948645
-
项目类别:Intergovernmental Personnel Award
-
资助金额:$21.51万
-
财政年份:2019
-
负责人:Amarda Shehu
-
依托单位:
Collaborative: SI2-SSE - A Plug-and-Play Software Platform of Robotics-Inspired Algorithms for Modeling Biomolecular Structures and Motions
-
批准号:1440581
-
项目类别:Standard Grant
-
资助金额:$21.73万
-
财政年份:2015
-
负责人:Amarda Shehu
-
依托单位:
Travel Awards for 2015 IEEE International Conference on Bioinformatics and Biomedicine (BIBM-2015)
-
批准号:1543744
-
项目类别:Standard Grant
-
资助金额:$2.18万
-
财政年份:2015
-
负责人:Amarda Shehu
-
依托单位:
CCF: AF: Small: Novel Stochastic Optimization Algorithms to Advance the Treatment of Dynamic Molecular Systems
-
批准号:1421001
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2014
-
负责人:Amarda Shehu
-
依托单位:
Workshop: 2014 NSF CISE CAREER Proposal Writing Workshop
-
批准号:1415210
-
项目类别:Standard Grant
-
资助金额:$7.38万
-
财政年份:2013
-
负责人:Amarda Shehu
-
依托单位:
CAREER: Probabilistic Methods for Addressing Complexity and Constraints in Protein Systems
-
批准号:1144106
-
项目类别:Continuing Grant
-
资助金额:$54.99万
-
财政年份:2012
-
负责人:Amarda Shehu
-
依托单位:
AF: Small: A Unified Computational Framework to Enhance the Ab-Initio Sampling of Native-Like Protein Conformations
-
批准号:1016995
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2010
-
负责人:Amarda Shehu
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: