课题基金 / 基金详情

MFB: Better Homologous Folding using Computational Linguistics and Deep Learning

MFB: Better Homologous Folding using Computational Linguistics and Deep Learning
MFB:使用计算语言学和深度学习更好的同源折叠
批准号:
2330737
负责人:
Liang Huang
金额:
$145.31万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-03-01 至 2027-02-28

项目摘要

项目成果

Liang Huang的其他基金

相似基金

相关文献

中文摘要
翻译
核糖核酸(RNA)在我们的日常生活中至关重要,因为它在每一个活细胞中都扮演着不可或缺的角色。此外,我们的世界最近被一种RNA病毒颠覆了,而这种病毒当时部分地被一种RNA疫苗所控制。与人们的普遍看法相反,RNA不仅仅是更广为人知的DNA和蛋白质之间的中间人,它还可以具有控制基因表达等深刻的生物学功能。这些功能是由RNA结构(RNA的“形状”)决定的,因此,对这些结构的准确建模对于理解RNA功能以及设计疫苗、检测试剂盒和药物至关重要。然而,现有的确定RNA结构的实验方法非常昂贵,而且往往局限于短序列,而且现有的计算工具相当慢,而且不完全准确。这种缓慢阻碍了它们在全长病毒基因组中的应用,如冠状病毒(约30,000个核苷酸或“字母”)。因此,迫切需要开发更好的计算方法来预测更准确、更有效和可扩展到更长序列(如整个基因组)的RNA结构。这一方向的进展可以提高我们对RNA病毒(包括普通感冒、流感、狂犬病、艾滋病毒、埃博拉病毒、脊髓灰质炎、麻疹等)的了解,并增强我们对抗下一次大流行的准备。该项目开发了预测多个相关(“同源”)RNA序列结构的有效算法,如SARS-CoV-2变种。这些算法将在平均序列长度和序列数量两者中线性扩展。这种线性缩放将使全基因组应用成为可能。研究人员的目标是利用人工智能(AI)的两个分支--自然语言处理和深度学习--的想法来实现这些目标。具体地说,本项目将改进三种类型的同源折叠算法,并将其应用于结构发现:(1)对齐-然后-折叠:首先对同源序列进行比对,然后预测对齐序列的一致结构;(2)迭代对齐-折叠:在序列比对和结构预测之间迭代;(3)同时对齐-折叠:联合预测比对和结构。该团队将采用这些快速方法,通过对RNA病毒基因组和转录本的全球结构预测来发现保守结构。这项研究将使发现新的RNA结构和功能成为可能,并将有助于疫苗、测试试剂盒和药物的设计。该项目得到了信息和智能系统部门、化学部门以及化学理论、模型和计算方法项目的支持。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Ribonucleic acid (RNA) is of utmost importance in our daily life because it plays essential roles in every living cell. Furthermore, our world was recently turned upside down by an RNA virus, which was then partially contained by an RNA vaccine. Contrary to common wisdom, RNA is not just an intermediate “messenger” between the more well-known DNA and protein, but it can also have profound biological functions such as controlling gene expression. These functions are determined by RNA structures (the “shapes” of the RNAs), and therefore accurate modeling of these structures is critical for understanding RNA functions and for designing vaccines, test kits, and drugs. However, existing experimental methods for determining RNA structure are extremely expensive and often limited to short sequences, and existing computational tools are rather slow and not completely accurate. This slowness hinders their applications to full-length viral genomes such as coronavirus (about 30,000 nucleotides or “letters”). Therefore, there is a critical need to develop better computational methods to predict RNA structures that are more accurate and more efficient and scalable to longer sequences such as whole genomes. Advances in this direction could improve our understanding of RNA viruses (which include common cold, influenza, Rabies, HIV, Ebola, polio, measles, and more) and increase our readiness to fight the next pandemic.This project develops efficient algorithms for predicting the structures of multiple related (“homologous”) RNA sequences such as SARS-CoV-2 variants. These algorithms will scale linearly in both the average sequence length and the number of sequences. This linear scaling will enable whole genome applications. The researchers aim to achieve these goals with ideas from two branches of artificial intelligence (AI): natural language processing and deep learning. Specifically, this project will improve three types of homologous folding algorithms and adapt them to structure discovery: (1) align-then-fold: first align the homologous sequences and then predict the consensus structure for the aligned sequences; (2) iteratively align-and-fold: iterate between sequence alignment and structure prediction; and (3) simultaneous align-and-fold: jointly predict alignment and structures. The team will adapt these fast methods to discover conserved structures using global structure prediction for RNA viral genomes and transcripts. This research will make it possible to discover new RNA structures and functions, and will help the design of vaccines, test kits, and drugs.This project is supported by the Divisions of Information and Intelligent Systems and of Chemistry and the Chemical Theory, Models, and Computational Methods Program in the Division of Chemistry.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Low-Latency and High-Quality Simultaneous Translation
  • 批准号:
    2009071
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2020
  • 负责人:
    Liang Huang
  • 依托单位:
RI: Small: Fast and Accurate Natural Language Parsing and Generation by Marrying Deep Learning with Dynamic Programming
  • 批准号:
    1817231
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2018
  • 负责人:
    Liang Huang
  • 依托单位:
EAGER: Collaborative Research: Scaling Up Discriminative Learning for Natural Language Understanding and Translation
  • 批准号:
    1656051
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.04万
  • 财政年份:
    2015
  • 负责人:
    Liang Huang
  • 依托单位:
EAGER: Collaborative Research: Scaling Up Discriminative Learning for Natural Language Understanding and Translation
  • 批准号:
    1449278
  • 项目类别:
    Standard Grant
  • 资助金额:
    $13.54万
  • 财政年份:
    2014
  • 负责人:
    Liang Huang
  • 依托单位:
海外基金