Natural selection and algorithmic design of mRNA

Natural selection and algorithmic design of mRNA
复制标题

DOI:
10.1089/10665270360688101
复制
发表时间:
2003-01-01
影响因子:
1.7
通讯作者:
Skiena, S
Skiena, S
中科院分区:
生物学4区
文献类型:
--
作者:
Cohen, B;Skiena, S

文献摘要

被引文献

相似文献

信使 RNA (mRNA) 序列根据三联体密码充当蛋白质的模板,其中 RNA 中的 4(3)=64 个不同密码子(三个连续核苷酸碱基的序列)中的每一个要么终止转录,要么映射到构建蛋白质的 20 个不同氨基酸(或残基)之一。由于密码子多于残基,因此编码中存在固有的冗余。某些残基(例如,色氨酸)仅具有单个相应的密码子,而其他残基(例如,精氨酸)具有多达六个相应的密码子。这种自由意味着编码给定蛋白质的可能 RNA 序列的数量随着蛋白质的长度呈指数增长。因此,大自然有很大的自由度来选择信息相同但结构和能量上不同的 mRNA 序列。在本文中,我们探讨了大自然如何利用这种自由,以及如何通过算法设计比通过自然选择构建的结构更有利的结构。特别是:(1)自然选择——我们进行了第一个大规模计算实验,比较了来自多种生物体的 mRNA 序列与尊重生物体密码子偏好的随机同义序列的稳定性。该实验针对来自 34 个微生物物种、36 个基因组结构的 27,000 多个序列进行。我们提供的证据表明,在所有基因组结构中,高度稳定的序列异常丰富,并且在 36 个案例中的 19 个案例中,高度不稳定的序列异常丰富。这表明 mRNA 序列的稳定性受到自然选择的影响。 (2)人工选择——受这些生物学结果的启发,我们研究了设计编码目标蛋白的最稳定和不稳定的mRNA序列的算法问题。我们给出了最稳定序列问题(MSSP)的多项式时间动态规划解决方案,该解决方案渐近并不比二级结构预测更复杂。我们证明了相应的最不稳定序列问题(LSSP)是 NP 完全的,并开发了两种启发式方法来构造此类序列。我们已经实现了这些算法,并展示了将高/低稳定性序列置于野生型和随机编码环境中的实验结果。我们的实现已经应用于 RNA“码字”的设计,在 RNA 计算中几乎不创建二级结构(Brenneman 和 Condon,2001;Marathe 等人,2001),并且我们预计这项工作将在序列设计问题上有各种其他应用(Skiena,2001)。
Messenger RNA (mRNA) sequences serve as templates for proteins according to the triplet code, in which each of the 4(3)=64 different codons (sequences of three consecutive nucleotide bases) in RNA either terminate transcription or map to one of the 20 different amino acids (or residues) which build up proteins. Because there are more codons than residues, there is inherent redundancy in the coding. Certain residues (e.g., tryptophan) have only a single corresponding codon, while other residues (e.g., arginine) have as many as six corresponding codons. This freedom implies that the number of possible RNA sequences coding for a given protein grows exponentially in the length of the protein. Thus nature has wide latitude to select among mRNA sequences which are informationally equivalent, but structurally and energetically divergent. In this paper, we explore how nature takes advantage of this freedom and how to algorithmically design structures more energetically favorable than have been built through natural selection. In particular: (1) Natural Selection-we perform the first large-scale computational experiment comparing the stability of mRNA sequences from a variety of organisms to random synonymous sequences which respect the codon preferences of the organism. This experiment was conducted on over 27,000 sequences from 34 microbial species with 36 genomic structures. We provide evidence that in all genomic structures highly stable sequences are disproportionately abundant, and in 19 of 36 cases highly unstable sequences are disproportionately abundant. This suggests that the stability of mRNA sequences is subject to natural selection. (2) Artificial Selection-motivated by these biological results, we examine the algorithmic problem of designing the most stable and unstable mRNA sequences which code for a target protein. We give a polynomial-time dynamic programming solution to the most stable sequence problem (MSSP), which is asymptotically no more complex than secondary structure prediction. We show that the corresponding least stable sequence problem (LSSP) is NP-complete, and develop two heuristics for the construction of such sequences. We have implemented these algorithms, and present experimental results placing the high/low stability sequences in context with both wildtype and random encodings. Our implementation has already been applied to the design of RNA "code-words" creating little or no secondary structure in RNA computing (Brenneman and Condon, 2001; Marathe et al., 2001), and we anticipate a variety of other applications of this work to sequence design problems (Skiena, 2001).