AREA: Optimizing gene expression with mRNA free energy modeling and algorithms
AREA: Optimizing gene expression with mRNA free energy modeling and algorithms
批准号:
8689532
负责人:
DANIEL PAUL AALBERTS
金额:
$25.53万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-06-01 至 2018-05-31
关键词:
AdenineAlgorithm DesignAlgorithmsAmino Acid SequenceAmino AcidsBase CompositionBindingBiologyBiotechnologyCodeCodon NucleotidesCollaborationsCommunitiesDataData SetDependenceEngineeringFree EnergyGene ExpressionGene Expression RegulationGenesGoalsGuanineLearningLinkLiteratureMapsMessenger RNAMetricMicroRNAsMiningModelingMutateMutationNucleotidesPeptide Sequence DeterminationPilot ProjectsProcessProductionProtein OverexpressionProteinsRNA FoldingRNA SequencesRNA SplicingRecombinantsResearchRibosomesRoleSmall Interfering RNASmall RNASolubilityStructureSynthetic GenesSystemTranslationsVaccine ProductionVariantbasecomputer programcostdata miningdesigndisorder riskdrug discoverydrug productionenergy densityexpression vectorhuman diseaseimprovedmodel designnovelpolypeptideprotein expressionprotein structureprototypepublic health relevanceresearch studystructural biologystructural genomicstool
中文摘要
描述(申请人提供):蛋白质过度表达对于从疫苗生产到药物发现的许多生物技术应用是可取的。基于与东北结构基因组学会合作的约20,000个常用表达载体的基因表达数据,我们观察到编码前50个核苷酸的自由能对多肽的表达水平有很强的预测作用。许多mRNA序列编码完全相同的蛋白质序列,因为多个密码子可能映射到相同的氨基酸。最近出现的基因表达水平的实验数据集为通过建模和算法设计最大化蛋白质表达创造了机会。我们建议开发算法,使生物学家能够评估天然mRNA序列是否可能高表达,并构建旨在优化基因表达的同义mRNA序列。在编码区的起始处具有相对较少的mRNA二级结构特别重要,因为这是核糖体组装的地方。过多的基因二级结构似乎也是有害的。在翻译、剪接和小干扰RNA基因调控机制中,信使RNA的一个区域必须被展开,以允许核糖体、剪接因子或microRNAs的结合。了解正在展开的自由能源成本为理解生物学和通过算法设计基因表达变化提供了机会。
英文摘要
DESCRIPTION (provided by applicant): Protein overexpression is desirable for many biotechnology applications ranging from vaccine production to drug discovery. Based on gene expression data of about 20,000 genes with common expression vectors from collaboration with the Northeast Structural Genomics Consortium, we observe that the free energy of the first ~50 coding nucleotides is strongly predictive of the expression level of polypeptides. Many mRNA sequences encode exactly the same protein sequence because multiple codons may map to the same amino acid. The recent emergence of experimental datasets of expression levels for genes has created an opportunity to maximize protein expression through modeling and algorithm design. We propose to develop algorithms that will enable biologists to evaluate whether native mRNA sequences are likely to express highly and to build synonymous mRNA sequences designed to optimize gene expression. Having relatively little mRNA secondary structure at the start of the coding region is of particular importance as that is where the ribosome assembles. Extensive mRNA secondary structure later in the gene also appears to be deleterious. In translation, splicing, and small interfering RNA gene regulation mechanisms, a region of messenger RNA must be unfolded to allow binding of the ribosome, splice factors, or microRNAs. Understanding the unfolding free energy costs offers opportunities to understand the biology of and to algorithmically engineer changes in gene expression.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Binding and Splicing mRNA
-
批准号:7924463
-
项目类别:
-
资助金额:$7.92万
-
财政年份:2009
-
负责人:DANIEL PAUL AALBERTS
-
依托单位:
Binding and Splicing mRNA
-
批准号:7252941
-
项目类别:
-
资助金额:$22.05万
-
财政年份:2007
-
负责人:DANIEL PAUL AALBERTS
-
依托单位:
Splicing, Folding, and Stretching Nucleic Acids
-
批准号:6666541
-
项目类别:
-
资助金额:$15.52万
-
财政年份:2003
-
负责人:DANIEL PAUL AALBERTS
-
依托单位:
海外基金