Interpretable and extendable deep learning model for biological sequence analysis and prediction
Interpretable and extendable deep learning model for biological sequence analysis and prediction
批准号:
10409152
负责人:
DONG XU
金额:
$23.48万
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-05-01 至 2023-04-30
关键词:
ATAC-seqAddressAlgorithmic SoftwareAlzheimer&aposs DiseaseAttentionBenchmarkingBioinformaticsBiologicalBiological ModelsBiologyBiotechnologyCell CommunicationCellsCodeCollaborationsCommunicationCommunitiesComplex AnalysisComputer AnalysisComputersDataData AnalysesData SetDatabasesDevelopmentDimensionsDiseaseEnvironmentEvaluationFormulationGenesGenomicsGraphHeadHeterogeneityIndividualKnowledgeMachine LearningMalignant NeoplasmsMeasuresMedicineMethodsModelingMultiomic DataNatureOhioPerformancePlayProblem FormulationsProcessPublic HealthPublishingRegulator GenesReportingResearchResearch PersonnelResourcesRoleRunningSequence AnalysisSiteSourceStructureSystemTechniquesTechnologyTestingTrainingUniversitiesValidationVisualizationWorkanalysis pipelinebasecell typedata complexitydata formatdata integrationdeep learningdeep learning algorithmexperienceflexibilityimprovedinnovationinsightlearning communitylearning strategymethod developmentneural networknovelonline resourceparent grantsingle cell sequencingsingle cell technologysingle-cell RNA sequencingtooltool developmenttranscriptomicsweb interfaceweb portalweb site
中文摘要
摘要
单细胞测序技术为研究生物学和医学提供了巨大的机会,但
计算分析往往是揭示生物学见解和定义细胞异质性的瓶颈
作为数据的基础。机器学习,特别是深度学习的应用前景广阔
应对这些挑战。虽然来自不同实验室的ML研究,包括PI的实验室,已经取得了显著的进展
在这方面的进展,ML社区在单细胞数据分析中的参与受到限制,因为
技术复杂性和生物学知识的障碍。为了吸引更多的ML专家进入这个领域,PI
提出了使大规模单细胞测序数据的ML准备就绪,并提供ML友好的开发
环境。具体目标包括:(1)收集、处理和管理各种单细胞测序数据
让他们做好ML准备。我们将从公共来源收集单细胞测序数据并将其转换为
转换成便于存储和处理的格式。数据将使用多个选项进行处理,例如
使用待开发的管道进行归算、归一化和降维。(2)配置数据
转化为基准。我们将使用收集的数据来构建基准,收集公共基准,以及
鼓励社区提交他们的基准。数据将分为培训、验证和
多个环境中的测试集,包括帮助高效方法开发的最低可行基准
以及全面评估的综合基准。我们将开发实用程序,以基于
一套评估措施,并生成详细的报告。我们将选择一组公共工具来运行它们
这些基准作为其他人进行比较的基线。(三)提供融合发展
支持部分方法开发的环境(IDE)。我们将构建一个用于单细胞测序的IDE
具有即插即用特性的分析方法开发在代码级别和用于ML的Web界面
研究人员贡献并测试任何最低限度的新想法。将提供一份包含评估的报告
计算机资源的度量和使用、与一些公共工具的比较以及下游可视化
和解释。新格式化的数据、基准以及方法开发和评估
GitHub和内部单细胞数据分析门户网站DeepMAPS将提供这一环境。这个
拟议的研究是母基金(R35-GM126985)的自然延伸,旨在开发深层次的
用于分析和预测生物序列的学习算法、工具和网络资源,包括(1)
开发通用的无监督表示法并使深度学习模型可解释
理解生物学机制和生成假说;(2)将深度学习模型应用于广泛的
生物信息学问题的范围,以及(3)使数据、模型和工具可供研究人员免费使用
社区。由于R35机制的灵活性,PI的实验室将这些方法扩展到单细胞
数据分析,这为实验室为拟议的任务做好了准备。
英文摘要
SUMMARY
Single-cell sequencing technologies provide great opportunities for studying biology and medicine, but
computational analyses are often the bottlenecks to reveal biological insights and define cellular heterogeneity
underlying the data. The applications of machine learning (ML), especially deep learning hold great promises
to address the challenges. While ML studies from various labs, including the PI’s lab, have made significant
progress along this line, the involvement of the ML community in single-cell data analysis is limited due to the
barriers of technology complexity and biology knowledge. To attract more ML experts into this field, the PI
proposes to make large-scale single-cell sequencing data ML-ready and provide an ML-friendly development
environment. Specific aims include: (1) Collect, process, and manage diverse single-cell sequencing data
to make them ML-ready. We will collect single-cell sequencing data from public sources and convert them
into formats efficient for storage and handling. The data will be processed with multiple options, such as
imputation, normalization, and dimension reduction using a pipeline to be developed. (2) Configure the data
into benchmarks. We will use the collected data to build benchmarks, gather public benchmarks, and
encourage the community to submit their benchmarks. The data will be divided into training, validation, and
test sets in multiple settings, including a minimum viable benchmark to assist efficient method development
and a comprehensive benchmark for full evaluations. We will develop utilities to evaluate results based on a
set of assessment measures, and generate detailed reports. We will select a set of public tools to run them on
the benchmarks as baselines for others to compare with. (3) Provide an integrated development
environment (IDE) to support partial method development. We will build an IDE for single-cell sequencing
analysis method development with plug-and-play features at the code level and web interface for ML
researchers to contribute and test any minimum new ideas. A report will be provided containing evaluation
metrics and usage of computer resources, comparisons with some public tools, and downstream visualization
and interpretation. The newly formatted data, the benchmarks, and the method development and assessment
environment will be available at GitHub and the in-house single-cell data analysis web portal DeepMAPS. The
proposed research is a natural extension of the parent grant (R35-GM126985), which aims to develop deep-
learning algorithms, tools, web resources for analyses and predictions of biological sequences, including (1)
developing general unsupervised representations and making deep-learning models interpretable for
understanding biological mechanisms and generating hypotheses; (2) applying deep-learning models to a wide
range of bioinformatics problems, and (3) making the data, models, and tools freely accessible to the research
community. Thanks to the flexibility of the R35 mechanism, the PI’s lab extended these methods to single-cell
data analyses, which well-prepared the lab for the proposed tasks.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Multi-view self-supervised deep learning for biological sequences and beyond
-
批准号:10623063
-
项目类别:
-
资助金额:$39.13万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
-
批准号:10395451
-
项目类别:
-
资助金额:$45.64万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
-
批准号:9925232
-
项目类别:
-
资助金额:$37.82万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Deep learning for protein subcellular/sub-organelle localizations and localization motifs
-
批准号:9768571
-
项目类别:
-
资助金额:$20.53万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8656715
-
项目类别:
-
资助金额:$27.89万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8258610
-
项目类别:
-
资助金额:$27.94万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8469528
-
项目类别:
-
资助金额:$26.94万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:9086384
-
项目类别:
-
资助金额:$27.84万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7648313
-
项目类别:
-
资助金额:$21.87万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7267931
-
项目类别:
-
资助金额:$13.79万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7881473
-
项目类别:
-
资助金额:$21.97万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7651361
-
项目类别:
-
资助金额:$22.03万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7138874
-
项目类别:
-
资助金额:$14.23万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
海外基金