Interpretable and extendable deep learning model for biological sequence analysis and prediction
Interpretable and extendable deep learning model for biological sequence analysis and prediction
批准号:
10395451
负责人:
DONG XU
金额:
$45.64万
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-05-01 至 2023-07-31
关键词:
Algorithmic SoftwareAmino Acid SequenceAreaBase SequenceBig DataBioinformaticsBiologicalBiological ModelsBiologyBiomedical ResearchCommunitiesComputational BiologyComputational algorithmDNADNA SequenceDataData AnalysesDevelopmentGenotypeGoalsHealthcareInformation SystemsKnowledgeLabelLearningLightMachine LearningMalignant NeoplasmsMedicalMedicineMethodsMicrobeModelingMutationMutation AnalysisPaperPerformancePhenotypePlantsPlug-inPost-Translational Protein ProcessingPropertyProteinsPublic HealthPublishingRNARNA SequencesResearchResource InformaticsSequence AnalysisSeriesSourceSystemTechnologyWorkcomputerized toolsdeep learningdeep learning algorithmdeep learning modeldesigndrug developmentimprovedin silicoindexingintegrated circuitlearning strategymachine learning methodmobile applicationnovelonline resourceopen sourcepersonalized diagnosticspersonalized medicineprecision medicineprotein structure functionprotein structure predictionsoftware systemssupervised learningsynthetic biologytoolunsupervised learning
中文摘要
项目摘要
生物信息学和计算生物学已成为生物医学研究的核心。少年派董旭博士
这一领域的工作重点是开发新的计算算法、软件和信息
系统以及这些工具和其他信息学资源在不同生物领域的广泛应用
以及医疗问题。他致力于蛋白质结构预测、翻译后研究的许多问题
在植物、微生物和微生物的计算机研究中的修饰预测、高通量生物数据分析
癌症、生物信息系统和医疗保健移动应用程序开发。他出版了更多
超过300篇论文,被引用约12000篇,H指数为55。在这个项目中,PI建议开发
用于生物序列分析和预测的深度学习算法、工具、网络资源,包括
DNA、RNA和蛋白质序列。这些数据的可获得性为精确度提供了新的机会
医学等领域,而深度学习作为机器学习中的一项前沿技术,呈现出一种
分析和预测生物序列的新的强有力的方法。随着快速积累
序列数据和深度学习方法的快速发展,迫切需要系统化
研究如何在序列分析和预测中最好地应用深度学习。为此目的,PI将
制定尖端深度学习方法,未来五年的目标如下:
(1)开发一系列新的深度学习方法和模型,专门针对生物
序列分析和预测:(A)DNA/RNA、蛋白质和
SNP/突变序列,为各种应用捕捉局部和全局特征;(B)方法
使深度学习模型可用于理解生物机制和生成
假设;(C)“规则学习”,通过结合无监督的学习来抽象潜在的“规则”
对大的无标签数据和小的有标签的数据进行有监督的学习,从而对新的无标签数据进行分类。
(2)将所提出的深度学习模型应用于dna/rna序列标注、基因-表型。
分析,癌症突变分析,蛋白质功能/结构预测,蛋白质定位预测,以及
蛋白质翻译后修饰预测。PI将利用与每个组件相关联的特定属性
针对这些问题,改进深度学习模型。他将制定一套相关的预测和分析
工具,这将改善最先进的性能,并揭示一些相关的生物学机制。
(3)向研究界免费提供数据、模型和工具。该系统将是
设计的模块化和开源,可通过GitHub获得。它们将像集成电路一样面市
模块,这些模块是通用的,可以插入到不同的应用中。PI将开发一个Web资源
有关生物序列的表示、分析和预测,以及帮助生物学家
没有计算知识来将深度学习应用于他们特定的研究问题。
英文摘要
Project Abstract
Bioinformatics and computational biology have become the core of biomedical research. The PI Dr. Dong Xu's
work in this area focuses on development of novel computational algorithms, software and information
systems, as well as on broad applications of these tools and other informatics resources for diverse biological
and medical problems. He works on many research problems in protein structure prediction, post-translational
modification prediction, high-throughput biological data analyses, in silico studies of plants, microbes and
cancers, biological information systems, and mobile App development for healthcare. He has published more
than 300 papers, with about 12,000 citations and H-index of 55. In this project, the PI proposes to develop
deep-learning algorithms, tools, web resources for analyses and predictions of biological sequences, including
DNA, RNA, and protein sequences. The availability of these data provides emerging opportunities for precision
medicine and other areas, while deep learning as a cutting-edge technology in machine learning, presents a
new powerful method for analyses and predictions of biological sequences. With rapidly accumulating
sequence data and fast development of deep-learning methods, there is an urgent need to systematically
investigate how to best apply deep learning in sequence analyses and predictions. For this purpose, the PI will
develop cutting-edge deep-learning methods with the following goals for the next five years:
(1) Develop a series of novel deep-learning methods and models to specifically target biological
sequence analyses and predictions in: (a) general unsupervised representations of DNA/RNA, protein and
SNP/mutation sequences that capture both local and global features for various applications; (b) methods to
make deep-learning models interpretable for understanding biological mechanisms and generating
hypotheses; (c) “rule learning”, which abstracts the underlying “rules” by combining unsupervised learning of
large unlabeled data and supervised learning of small labeled data so that it can classify new unlabeled data.
(2) Apply the proposed deep-learning model to DNA/RNA sequence annotation, genotype-phenotype
analyses, cancer mutation analyses, protein function/structure prediction, protein localization prediction, and
protein post-translational modification prediction. The PI will exploit particular properties associated with each
of these problems to improve the deep-learning models. He will develop a set of related prediction and analysis
tools, which will improve the state-of-art performance and shed some light on related biological mechanisms.
(3) Make the data, models, and tools freely accessible to the research community. The system will be
designed modular and open-source, available through GitHub. They will be available like integrated circuit
modules, which are universal and ready to plug in for different applications. The PI will develop a web resource
for biological sequence representations, analyses, and predictions, as well as tutorials to help biologists with
no computational knowledge to apply deep learning to their specific research problems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Multi-view self-supervised deep learning for biological sequences and beyond
-
批准号:10623063
-
项目类别:
-
资助金额:$39.13万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
-
批准号:9925232
-
项目类别:
-
资助金额:$37.82万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Deep learning for protein subcellular/sub-organelle localizations and localization motifs
-
批准号:9768571
-
项目类别:
-
资助金额:$20.53万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
-
批准号:10409152
-
项目类别:
-
资助金额:$23.48万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8656715
-
项目类别:
-
资助金额:$27.89万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8258610
-
项目类别:
-
资助金额:$27.94万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8469528
-
项目类别:
-
资助金额:$26.94万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:9086384
-
项目类别:
-
资助金额:$27.84万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7648313
-
项目类别:
-
资助金额:$21.87万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7267931
-
项目类别:
-
资助金额:$13.79万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7881473
-
项目类别:
-
资助金额:$21.97万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7651361
-
项目类别:
-
资助金额:$22.03万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7138874
-
项目类别:
-
资助金额:$14.23万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
海外基金