Interpretable and extendable deep learning model for biological sequence analysis and prediction
Interpretable and extendable deep learning model for biological sequence analysis and prediction
批准号:
10409152
负责人:
DONG XU
金额:
$23.48万
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-05-01 至 2023-04-30
关键词:
ATAC-seqAddressAlgorithmic SoftwareAlzheimer&aposs DiseaseAttentionBenchmarkingBioinformaticsBiologicalBiological ModelsBiologyBiotechnologyCell CommunicationCellsCodeCollaborationsCommunicationCommunitiesComplex AnalysisComputer AnalysisComputersDataData AnalysesData SetDatabasesDevelopmentDimensionsDiseaseEnvironmentEvaluationFormulationGenesGenomicsGraphHeadHeterogeneityIndividualKnowledgeMachine LearningMalignant NeoplasmsMeasuresMedicineMethodsModelingMultiomic DataNatureOhioPerformancePlayProblem FormulationsProcessPublic HealthPublishingRegulator GenesReportingResearchResearch PersonnelResourcesRoleRunningSequence AnalysisSiteSourceStructureSystemTechniquesTechnologyTestingTrainingUniversitiesValidationVisualizationWorkanalysis pipelinebasecell typedata complexitydata formatdata integrationdeep learningdeep learning algorithmexperienceflexibilityimprovedinnovationinsightlearning communitylearning strategymethod developmentneural networknovelonline resourceparent grantsingle cell sequencingsingle cell technologysingle-cell RNA sequencingtooltool developmenttranscriptomicsweb interfaceweb portalweb site
中文摘要
总结
英文摘要
SUMMARY
Single-cell sequencing technologies provide great opportunities for studying biology and medicine, but
computational analyses are often the bottlenecks to reveal biological insights and define cellular heterogeneity
underlying the data. The applications of machine learning (ML), especially deep learning hold great promises
to address the challenges. While ML studies from various labs, including the PI’s lab, have made significant
progress along this line, the involvement of the ML community in single-cell data analysis is limited due to the
barriers of technology complexity and biology knowledge. To attract more ML experts into this field, the PI
proposes to make large-scale single-cell sequencing data ML-ready and provide an ML-friendly development
environment. Specific aims include: (1) Collect, process, and manage diverse single-cell sequencing data
to make them ML-ready. We will collect single-cell sequencing data from public sources and convert them
into formats efficient for storage and handling. The data will be processed with multiple options, such as
imputation, normalization, and dimension reduction using a pipeline to be developed. (2) Configure the data
into benchmarks. We will use the collected data to build benchmarks, gather public benchmarks, and
encourage the community to submit their benchmarks. The data will be divided into training, validation, and
test sets in multiple settings, including a minimum viable benchmark to assist efficient method development
and a comprehensive benchmark for full evaluations. We will develop utilities to evaluate results based on a
set of assessment measures, and generate detailed reports. We will select a set of public tools to run them on
the benchmarks as baselines for others to compare with. (3) Provide an integrated development
environment (IDE) to support partial method development. We will build an IDE for single-cell sequencing
analysis method development with plug-and-play features at the code level and web interface for ML
researchers to contribute and test any minimum new ideas. A report will be provided containing evaluation
metrics and usage of computer resources, comparisons with some public tools, and downstream visualization
and interpretation. The newly formatted data, the benchmarks, and the method development and assessment
environment will be available at GitHub and the in-house single-cell data analysis web portal DeepMAPS. The
proposed research is a natural extension of the parent grant (R35-GM126985), which aims to develop deep-
learning algorithms, tools, web resources for analyses and predictions of biological sequences, including (1)
developing general unsupervised representations and making deep-learning models interpretable for
understanding biological mechanisms and generating hypotheses; (2) applying deep-learning models to a wide
range of bioinformatics problems, and (3) making the data, models, and tools freely accessible to the research
community. Thanks to the flexibility of the R35 mechanism, the PI’s lab extended these methods to single-cell
data analyses, which well-prepared the lab for the proposed tasks.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Multi-view self-supervised deep learning for biological sequences and beyond
-
批准号:10623063
-
项目类别:
-
资助金额:$39.13万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
-
批准号:10395451
-
项目类别:
-
资助金额:$45.64万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
-
批准号:9925232
-
项目类别:
-
资助金额:$37.82万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Deep learning for protein subcellular/sub-organelle localizations and localization motifs
-
批准号:9768571
-
项目类别:
-
资助金额:$20.53万
-
财政年份:2018
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8656715
-
项目类别:
-
资助金额:$27.89万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8258610
-
项目类别:
-
资助金额:$27.94万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:8469528
-
项目类别:
-
资助金额:$26.94万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
Development of MUFOLD for Building High-Accuracy Protein Structure Models
-
批准号:9086384
-
项目类别:
-
资助金额:$27.84万
-
财政年份:2012
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7648313
-
项目类别:
-
资助金额:$21.87万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7267931
-
项目类别:
-
资助金额:$13.79万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7881473
-
项目类别:
-
资助金额:$21.97万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7651361
-
项目类别:
-
资助金额:$22.03万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
New Scoring, Assembly and Evaulation Techiniques for Protein Structure Prediction
-
批准号:7138874
-
项目类别:
-
资助金额:$14.23万
-
财政年份:2006
-
负责人:DONG XU
-
依托单位:
海外基金