Predictive modeling of plant messenger RNA polyadenylation sites.

Predictive modeling of plant messenger RNA polyadenylation sites.
复制标题

植物信使RNA聚腺苷酸位点的预测建模。

DOI:
10.1186/1471-2105-8-43
复制
发表时间:
2007-02-07
期刊:
影响因子:
3
通讯作者:
Li, Qingshun Quinn
Li, Qingshun Quinn
中科院分区:
生物学4区
文献类型:
--
作者:
Ji, Guoli;Zheng, Jianti;Shen, Yingjia;Wu, Xiaohui;Jiang, Ronghan;Lin, Yun;Loke, Johnny C;Davis, Kimberly M;Reese, Greg J;Li, Qingshun Quinn

文献摘要

被引文献

相似文献

在前mRNA成熟过程中的一个重要的加工事件是多聚腺嘌呤[PolyA]尾巴的转录后添加。3‘端的聚(A)轨迹保护mRNA不受调控地降解,并通过mRNA输出和翻译机制的识别来指示mRNA的完整性。多聚(A)位点的位置由多聚腺苷酸化因子复合体识别的前mRNA序列中的信号预先确定。这些信号通常是作为未来聚(A)位点的裂解位点周围的三部分序列模式。在植物中,这些信号元件之间几乎没有序列保守性,这使得开发一种准确的算法来预测给定基因的Poly(A)位点变得困难。我们试图解决这个问题。基于我们目前的工作模型和拟南芥Poly(A)信号和Poly(A)位点周围的核苷酸序列分布,我们设计了一个基于广义隐马尔可夫模型的算法来预测潜在的PolyA位点。通过对多个数据集的测试,证明了该算法具有很高的特异度和灵敏度,在最佳组合下,两者都达到了97%。该程序的准确性,称为聚(A)位点侦察或通过,已经证明了许多有效的聚(A)位点的预测。PASS还预测了通过传统遗传实验构建和表征的聚(A)信号突变体中聚(A)位点效率的变化。通过预测长基因组序列中的Poly(A)位点,证明了PASS的有效性。根据植物Poly(A)信号的特点,建立了一个有效预测拟南芥基因PolyA位点的计算模型。该算法将在基因注释中有用,因为聚(A)位点表示转录本的结束。该算法还可用于预测已知基因中的替代多聚(A)位点,通过预测和消除不需要的多聚(A)位点,将在作物基因工程的转基因设计中有用。
One of the essential processing events during pre-mRNA maturation is the post-transcriptional addition of a polyadenine [poly(A)] tail. The 3'-end poly(A) track protects mRNA from unregulated degradation, and indicates the integrity of mRNA through recognition by mRNA export and translation machinery. The position of a poly(A) site is predetermined by signals in the pre-mRNA sequence that are recognized by a complex of polyadenylation factors. These signals are generally tri-part sequence patterns around the cleavage site that serves as the future poly(A) site. In plants, there is little sequence conservation among these signal elements, which makes it difficult to develop an accurate algorithm to predict the poly(A) site of a given gene. We attempted to solve this problem. Based on our current working model and the profile of nucleotide sequence distribution of the poly(A) signals and around poly(A) sites in Arabidopsis, we have devised a Generalized Hidden Markov Model based algorithm to predict potential poly(A) sites. The high specificity and sensitivity of the algorithm were demonstrated by testing several datasets, and at the best combinations, both reach 97%. The accuracy of the program, called poly(A) site sleuth or PASS, has been demonstrated by the prediction of many validated poly(A) sites. PASS also predicted the changes of poly(A) site efficiency in poly(A) signal mutants that were constructed and characterized by traditional genetic experiments. The efficacy of PASS was demonstrated by predicting poly(A) sites within long genomic sequences. Based on the features of plant poly(A) signals, a computational model was built to effectively predict the poly(A) sites in Arabidopsis genes. The algorithm will be useful in gene annotation because a poly(A) site signifies the end of the transcript. This algorithm can also be used to predict alternative poly(A) sites in known genes, and will be useful in the design of transgenes for crop genetic engineering by predicting and eliminating undesirable poly(A) sites.