Interpretable Computational Models of Functional Genomics Data
Interpretable Computational Models of Functional Genomics Data
批准号:
10698090
负责人:
Peter K Koo
金额:
$43.2万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-07 至 2027-06-30
关键词:
5&apos Splice SiteAccelerationAddressAlternative SplicingBayesian AnalysisBiologicalBiological AssayBiological ProcessBiological SciencesCatalogsChIP-seqCodeCommunitiesComplexComputer ModelsComputer softwareComputing MethodologiesDNADNA SequenceDataDependenceDevelopmentDiseaseElementsExhibitsGeneticGenetic TranscriptionGenomeGenomicsGoalsIndividualKnowledgeLaboratoriesLearningMapsMethodsModelingModernizationMutagenesisMutagensNetwork-basedNucleotidesPerformancePositioning AttributeRNARNA-Binding ProteinsRegulatory ElementResolutionSpecific qualifier valueTrainingTranscriptional RegulationTranslatingVariantWeightWorkbasecomputerized toolsconvolutional neural networkcrosslinking and immunoprecipitation sequencingdensitydesigndirect applicationexperimental studyfunctional genomicsgenome-widegenomic datagenomic locushuman diseaseimprovedin silicoin vivoinsightlearning networkmachine learning methodmultiplex assayneural network architectureopen sourceprototypesequence learningsyntaxtranscription factoruser-friendlyweb server
中文摘要
点击翻译按钮获取中文摘要
英文摘要
PROJECT SUMMARY
Understanding how the coordination of cis-regulatory elements (CREs) influences biological processes, such as
transcription and alternative splicing, is a major goal in computational genomics. This remains a challenge
because CRE activity at any given locus may depend on a host of other factors, including sequence context
and/or the presence of other CREs nearby. Recent developments in deep convolutional neural networks (CNNs)
have revolutionized our ability to predict regulatory functions from DNA sequence. Unlike previous computational
methods based on position-weight matrices, which capture an additive model of CREs, CNNs can, in principle,
also learn higher-order dependencies within the CRE, with other CREs, and with the broader sequence context.
However, CNNs are essentially black box models, with parameters that don’t have clear biological meaning.
Hence it remains a challenge to translate the improved predictions of a CNN to new biological insights. Here we
propose to develop three different computational methods that can comprehensively characterize higher-order
interactions within CREs and across different CREs from functional genomics data, specifically ChIP-seq and
CLIP-seq data publicly available through ENCODE. Each method serves as its own separate Aim and will be
developed in parallel. In Aim 1, we will develop a new post hoc model interpretability method based on employing
interpretable quantitative models originally developed to understand complex genetic interactions in laboratory-
based comprehensive mutagenesis (e.g. multiplex assays of variant effects) to characterize CRE dependencies
learned by a CNN, using synthetic sequences to target specific biological hypotheses. In Aim 2, we will develop
new CNN architectures where the learned parameters will express higher-order interactions that have direct
biological interpretations. In Aim 3, we will combine a Bayesian nonparametric framework for modeling CREs
with CNN-based CRE annotations and GPU acceleration to develop new methods for understanding how CREs
are specified in the genome. Successful completion of these Aims will provide a leap forward in our
understanding of higher-order CRE dependencies that are exploited but have not yet been fully revealed by
CNNs. This work will provide the community with: (1) a new suite of open-source computational tools that
address the problem of modeling CREs and their dependencies in functional genomics data; and (2) a
comprehensive genome-wide catalogue of CRE syntax for transcription factors and RNA-binding proteins that
will be hosted on a user-friendly webserver.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Reliable post hoc interpretations of deep learning in genomics
-
批准号:10638753
-
项目类别:
-
资助金额:$38.4万
-
财政年份:2023
-
负责人:Peter K Koo
-
依托单位:
Interpretable Computational Models of Functional Genomics Data
-
批准号:10453055
-
项目类别:
-
资助金额:$41.73万
-
财政年份:2022
-
负责人:Peter K Koo
-
依托单位:
海外基金