A predictive model of mRNA stability and translation for variant interpretation and mRNA therapeutics
A predictive model of mRNA stability and translation for variant interpretation and mRNA therapeutics
批准号:
9894822
负责人:
Georg Seelig
金额:
$47.31万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-05 至 2021-03-31
关键词:
3&apos Untranslated Regions5&apos Untranslated RegionsAddressAffectAlternative SplicingBinding SitesBiologicalBiological AssayBiologyBiotechnologyCodeComputational TechniqueDNA LibraryDataData SetDiseaseElementsEngineeringExpression LibraryGene ExpressionGenesGenetic ProgrammingGenetic TranscriptionGenetic TranslationGenetic VariationGenomeGenomicsHumanHuman EngineeringHuman GeneticsHuman GenomeImageIn VitroLearningLibrariesMachine LearningMeasuresMessenger RNAMethodsModelingNeural Network SimulationPerformancePolyribosomesProductionPropertyProteinsRNARNA SplicingRNA StabilityRNA-Binding ProteinsRandomizedRegulator GenesRegulatory ElementResearch Project GrantsResolutionRibosomesRoleSiteSourceStructureTechniquesTestingTherapeuticThiouridineTimeTrainingTranscriptTranslatingTranslationsUntranslated RegionsValidationVariantWorkbasecomparativecomputer scienceconvolutional neural networkdesignexperimental studygenetic variantmRNA Stabilitymachine visionmemberneural networknovelpolysome profilingpractical applicationpredictive modelingprotein expressionribosome profilingscreeningstable cell linestatistical learningsynthetic constructvoice recognition
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The leading and trailing untranslated regions (UTRs) of an mRNA, along with the coding sequence (CDS),
control protein production by modulating translation and mRNA stability. However, although we have identified
a vast number of regulatory features in these regions, we are still far from being able to predict, for example,
whether and how a sequence variant affects the levels of protein being made. Here, we propose to combine
high-throughput experimental characterization of protein expression in synthetic libraries with machine learning
to create predictive models of translation and mRNA stability, addressing an urgent need. Recent progress in
machine vision, voice recognition and other fields of computer science has been driven by the availability of
enormous data sets on which to train models. Machine learning approaches have also had remarkable impact
in biology, but biological data sets often are comparatively small, limiting the quality of models that can be
learned. For example, there are only around 20,000 genes in the human genome, a restrictively small set of
examples for training a predictive model that captures the full extent of the genome’s “regulatory code.” In this
proposal, we aim to overcome this data size limitation by training predictive models of protein expression on
data from millions of synthetic constructs -- a data set several orders of magnitude larger than the number of
genes in the genome. Specifically, we will create libraries of in vitro transcribed mRNA with targeted variation
in the UTRs and CDS and will assay protein expression of each library member by performing high-throughput
polysome profiling, ribosome profiling, and mRNA stability assays. We will then use neural network
approaches to learn predictive models of the relationship between mRNA sequence and levels of protein
production. We will apply our models to three applications of practical importance: first, we expect to uncover
novel biology, for example identifying regulatory sequence elements and interactions between them. Second,
we will validate our models through the de novo design and experimental testing of sequences that result in
higher levels or protein production than any of the millions of randomly generated members of the original
library or than the endogenous UTR sequences currently used in biotechnology. Such stable and highly
translating mRNA constructs would be of particular value for the field or mRNA therapeutics. Third, we will
predict the functional consequences of genetic variation in UTRs on protein production and we will validate
these predictions experimentally. We are far from understanding which genetic variants compromise gene
regulatory function in ways that may contribute to disease, making such a comprehensive and quantitative
analysis of variants valuable.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Engineering cell type-specific splicing regulation
-
批准号:10633765
-
项目类别:
-
资助金额:$39.57万
-
财政年份:2023
-
负责人:Georg Seelig
-
依托单位:
Joint receptor and protein expression immunophenotyping through split-pool barcoding
-
批准号:10625987
-
项目类别:
-
资助金额:$39.62万
-
财政年份:2021
-
负责人:Georg Seelig
-
依托单位:
Joint receptor and protein expression immunophenotyping through split-pool barcoding
-
批准号:10375354
-
项目类别:
-
资助金额:$40.09万
-
财政年份:2021
-
负责人:Georg Seelig
-
依托单位:
High-resolution spatial transcriptomics through light patterning
-
批准号:9886581
-
项目类别:
-
资助金额:$21.81万
-
财政年份:2020
-
负责人:Georg Seelig
-
依托单位:
High-resolution spatial transcriptomics through light patterning
-
批准号:10341212
-
项目类别:
-
资助金额:$17.81万
-
财政年份:2020
-
负责人:Georg Seelig
-
依托单位:
A massively parallel reporter assay for measuring chromatin effects on alternative splicing
-
批准号:10161803
-
项目类别:
-
资助金额:$22.6万
-
财政年份:2020
-
负责人:Georg Seelig
-
依托单位:
A massively parallel reporter assay for measuring chromatin effects on alternative splicing
-
批准号:9977420
-
项目类别:
-
资助金额:$18.75万
-
财政年份:2020
-
负责人:Georg Seelig
-
依托单位:
High-resolution spatial transcriptomics through light patterning
-
批准号:10112854
-
项目类别:
-
资助金额:$18.17万
-
财政年份:2020
-
负责人:Georg Seelig
-
依托单位:
Predictive Modeling of Alternative Splicing and Polyadenylation from Millions of Random Sequences
-
批准号:9306648
-
项目类别:
-
资助金额:$59.66万
-
财政年份:2017
-
负责人:Georg Seelig
-
依托单位:
海外基金