Multi-view self-supervised deep learning for biological sequences and beyond
Multi-view self-supervised deep learning for biological sequences and beyond
批准号:
10623063
负责人:
DONG XU
金额:
$39.13万
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
未结题
起止时间:
2018-05-01 至 2028-07-31
关键词:
3-DimensionalAddressAmino Acid SequenceBase SequenceBioinformaticsBiologicalBiologyBiomedical ResearchCellsClustered Regularly Interspaced Short Palindromic RepeatsCommunitiesDataData AnalysesDevelopmentDiseaseDrug DesignGenesHumanInformation SystemsIonsJointsLabelLearningLigand BindingMainstreamingMedicalMethodsModalityModelingMolecular BiologyPaperPhylogenetic AnalysisProtein FamilyProteinsResearchSelf PerceptionSequence AnalysisShapesSoftware ToolsSourceSystemTechniquesTestingTrainingTreesartificial intelligence methoddeep learningdeep learning algorithmdesigndrug developmentlearning strategymolecular dynamicsmultimodal datanovelonline resourceopen source toolprotein structureprotein structure predictionsoftware systemssuccesssupervised learningtrend
中文摘要
项目摘要
深度学习(DL)在解决基本生物学问题方面的广度和深度一直受到关注。
演示。基于DL的方法,如用于3D蛋白质结构预测的AlphaFold 2,已经成为
被生物学界广泛接受。徐实验室一直处于开发新型DL的前沿
算法、软件和信息系统,用于解决各种生物和医学问题。于本
项目期间,徐实验室在解决一些紧迫的挑战和需求方面取得了出色的进展
用于开发生物序列分析和预测以及其他生物信息学中的DL方法
问题这个R35项目已经产生了31篇论文,涉及的研究主题从蛋白质序列-
基于预测的药物设计、分子动力学模拟和单细胞数据分析。此外还
还向社区提供了十多个开源工具和三个主要的网络资源。
新的DL技术的快速发展和徐实验室在该领域的专业知识积累带来了新的
在塑造DL分子生物学的机会。目前在生物医学中广泛使用的监督DL方法
研究往往没有足够的数据与干净和准确的标签训练,可能没有良好的
普遍性新兴的自我监督学习(SSL)方法旨在学习信息
通过暴露不同数据透视图之间的关系而无需人工注释来表示,
成为一种新趋势。不同的数据透视图被广泛地称为多视图。多视图SSL技术
使我们能够为单模态和多模态数据生成联合或协调表示,
概括性、更好的鲁棒性和更少的偏差。尽管SSL在其他领域取得了巨大的成功,
它在生物学中只被最低限度地探索过。
这个更新项目将开发一个多视图SSL框架,可以处理单视图和多视图。
查看数据并能够执行单个和多个任务。它将解决应用中的关键挑战和瓶颈
用于生物学研究的SSL,例如选择有效的视图和数据增强,融合多模态数据或
数据从异构来源,并集成生物约束到SSL模型。我们将专注于
设计一个生物学信息系统,提高概括性和鲁棒性,并使结果
生物学上可解释且置信度可评估。Xu实验室将应用并完善该框架,
主流生物学应用,包括抗CRISPR蛋白预测,通过探索各种数据
使用互补序列预测蛋白质序列、离子和小配体结合的增强方法
蛋白质序列和结构的视图,以及不同条件下的单细胞数据分析。的
框架也将在基于序列的研究及其他方面进行广泛应用的测试,例如比对-
免费的方法来构建系统发育树和检测新的蛋白质家族,以及进行
跨物种单细胞数据分析。
英文摘要
Project Abstract
The breadth and depth of deep learning (DL) in solving fundamental biological problems have been
demonstrated. DL-based approaches, such as AlphaFold2 for 3D protein structure prediction, have become
widely accepted by the biology community. The Xu lab has been at the forefront of developing novel DL
algorithms, software, and information systems for diverse biological and medical problems. During the current
project period, the Xu lab has made excellent progress in addressing some of the urgent challenges and needs
for developing DL methods in biological sequence analyses and predictions, as well as other bioinformatics
problems. This R35 project has produced 31 papers covering research topics ranging from protein sequence-
based predictions to drug design, molecular dynamics simulation, and single-cell data analysis. In addition, it
also provided more than ten open-source tools and three major web-based resources to the community.
The rapid development of new DL techniques and Xu lab’s accumulating expertise in this field bring new
opportunities in shaping DL to molecular biology. The current widely used supervised DL methods in biomedical
research often do not have sufficient data with clean and accurate labels for training and may not have good
generalizability. The emerging self-supervised learning (SSL) approaches that aim to learn informative
representations by exposing relationships between different data perspectives without human annotations are
becoming a new trend. Different data perspectives are broadly called multiview. The multi-view SSL techniques
allow us to generate joint or coordinated representations for single modal and multimodal data with stronger
generalizability, better robustness, and less bias. Though SSL has demonstrated great successes in other fields,
it has only been minimally explored in biology.
This renewal project will develop a multi-view SSL framework that can handle both single-view and multi-
view data and is capable of single and multiple tasks. It will tackle key challenges and bottlenecks in applying
SSL for biological studies, such as selecting effective views and data augmentations, fusing multimodal data or
data from heterogeneous sources, and integrating biological constraints into SSL models. We will focus on
designing a biology-informed system, enhancing generalizability and robustness, and making the results
biologically interpretable and confidence assessable. The Xu lab will apply and refine the framework to multiple
mainstream biology applications, including anti-CRISPR protein prediction, by exploring various data
augmentation methods for protein sequences, ion and small ligand binding prediction using complementary
views of protein sequences and structures, and single-cell data analyses across different conditions. The
framework will also be tested for broad applications in sequence-based studies and beyond, such as alignment-
free methods for constructing phylogenetic trees and detecting novel protein families, as well as conducting
cross-species single-cell data analysis.
期刊论文(22)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DeepDom: Predicting protein domain boundary from sequence alone using stacked bidirectional LSTM.
DeepDom:使用堆叠双向 LSTM 仅根据序列预测蛋白质域边界。
DOI:
--
发表时间:
2019
期刊:
Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
影响因子:
--
作者:
[Jiang,Yuexu, Wang,Duolin, Xu,Dong]
通讯作者:
Xu,Dong
DOI:
10.1093/nar/gkab407
发表时间:
2021-07-02
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Zeng S, Mao Z, Ren Y, Wang D, Xu D, Joshi T]
通讯作者:
Joshi T
DOI:
10.3390/molecules28196793
发表时间:
2023-09-25
期刊:
Molecules (Basel, Switzerland)
影响因子:
--
作者:
[Essien C, Jiang L, Wang D, Xu D]
通讯作者:
Xu D
DOI:
10.3389/fgene.2022.912813
发表时间:
2022
期刊:
Frontiers in genetics
影响因子:
3.7
作者:
[]
通讯作者:
DOI:
10.3389/fpls.2022.831204
发表时间:
2022
期刊:
Frontiers in plant science
影响因子:
5.6
作者:
[Su L, Xu C, Zeng S, Su L, Joshi T, Stacey G, Xu D]
通讯作者:
Xu D
共 12 条
Interpretable and extendable deep learning model for biological sequence analysis and prediction
-
批准号:10395451
-
项目类别:
-
资助金额:$45.64万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
-
批准号:9925232
-
项目类别:
-
资助金额:$37.82万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Deep learning for protein subcellular/sub-organelle localizations and localization motifs
-
批准号:9768571
-
项目类别:
-
资助金额:$20.53万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
-
批准号:10409152
-
项目类别:
-
资助金额:$23.48万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8656715
-
项目类别:
-
资助金额:$27.89万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8258610
-
项目类别:
-
资助金额:$27.94万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8469528
-
项目类别:
-
资助金额:$26.94万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:9086384
-
项目类别:
-
资助金额:$27.84万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7648313
-
项目类别:
-
资助金额:$21.87万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7267931
-
项目类别:
-
资助金额:$13.79万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7881473
-
项目类别:
-
资助金额:$21.97万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7651361
-
项目类别:
-
资助金额:$22.03万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7138874
-
项目类别:
-
资助金额:$14.23万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
海外基金