A multispecies polyadenylation site model.

A multispecies polyadenylation site model.
复制标题

DOI:
10.1186/1471-2105-14-s2-s9
复制
发表时间:
2013
期刊:
影响因子:
3
通讯作者:
Duffy S
Duffy S
中科院分区:
生物学4区
文献类型:
--
作者:
Ho ES;Gunderson SI;Duffy S

文献摘要

被引文献

相似文献

多聚腺苷酸化存在于生命的所有三个结构域中,使其成为与剪接和5 '-加帽相比最保守的转录后过程。尽管大多数哺乳动物多聚腺苷酸位点在上游区域含有高度保守的六核苷酸,在下游区域含有保守性低得多的富含U/GU的序列,但也有许多例外。此外,在其他物种,如植物和无脊椎动物中的多聚腺苷酸位点,表现出与这种基因组结构的高度偏差,使得构建一般的多聚腺苷酸位点识别模型具有挑战性。我们调查了1999年至2011年间发表的9种poly(A)位点预测方法。所有方法都利用跨多聚腺苷酸位点的偏斜核苷酸谱和高度保守的多聚腺苷酸信号作为识别的主要特征。这些方法通常使用大量的特征,这将模型的维度增加到严重的程度,并且通常不会针对许多种类的基因组进行验证。我们提出了一个多聚(A)网站模型,采用最小的功能来捕捉聚(A)网站的本质,但在不同的物种产生更好的预测精度。我们的模型包括三个二或三核苷酸的档案,通过主成分分析,和预测的多聚(A)位点两侧的核小体占有率。我们使用两种机器学习方法验证了我们的模型:逻辑回归和线性判别分析。结果表明,模型在7种动物和植物中达到85-92%的灵敏度和85-96%的特异性。当我们应用一个物种的模型来预测其他物种的poly(A)位点时,灵敏度评分与系统发育距离相关。一个面向小基序的四特征模型足以准确地学习和预测真核生物中的poly(A)位点。
Polyadenylation is present in all three domains of life, making it the most conserved post-transcriptional process compared with splicing and 5'-capping. Even though most mammalian poly(A) sites contain a highly conserved hexanucleotide in the upstream region and a far less conserved U/GU-rich sequence in the downstream region, there are many exceptions. Furthermore, poly(A) sites in other species, such as plants and invertebrates, exhibit high deviation from this genomic structure, making the construction of a general poly(A) site recognition model challenging. We surveyed nine poly(A) site prediction methods published between 1999 and 2011. All methods exploit the skewed nucleotide profile across the poly(A) sites, and the highly conserved poly(A) signal as the primary features for recognition. These methods typically use a large number of features, which increases the dimensionality of the models to crippling degrees, and typically are not validated against many kinds of genomes. We propose a poly(A) site model that employs minimal features to capture the essence of poly(A) sites, and yet, produces better prediction accuracy across diverse species. Our model consists of three dior-trinucleotide profiles identified through principle component analysis, and the predicted nucleosome occupancy flanking the poly(A) sites. We validated our model using two machine learning methods: logistic regression and linear discriminant analysis. Results show that models achieve 85-92% sensitivity and 85-96% specificity in seven animals and plants. When we applied one model from one species to predict poly(A) sites from other species, the sensitivity scores correlate with phylogenetic distances. A four-feature model geared towards small motifs was sufficient to accurately learn and predict poly(A) sites across eukaryotes.