Sequence determinants in human polyadenylation site selection.

Sequence determinants in human polyadenylation site selection.
复制标题

DOI:
10.1186/1471-2164-4-7
复制
发表时间:
2003-02-25
期刊:
影响因子:
4.4
通讯作者:
Gautheret D
Gautheret D
中科院分区:
生物学2区
文献类型:
--
作者:
Legendre M;Gautheret D

文献摘要

参考文献

被引文献

相似文献

差异聚腺苷酸化是高等真核生物在不同环境中产生不同3‘端mRNAs的一种普遍机制。这涉及到3‘非编码区中的几个可供选择的多聚腺苷酸化位点,每个位点都有其特定的强度。在这里,我们分析人类多聚腺苷信号附近的模式,以帮助区分强和弱多聚腺苷酸化位点,或从随机出现的信号中区分真实位点。我们使用人类基因组序列来检索多聚腺苷化信号下游的区域,该区域通常在cdna或mrna数据库中缺失。分析了4956个EST验证的多聚腺苷酸化位点及其-300/+300nt侧翼区,我们清楚地显示了上游(Use)和下游(DSE)序列元件,它们都具有富含U(而不是富Gu)片段的特征。USE和DSE的存在是区分真正的多腺苷基化位点和随机出现的A(A/U)UAAA六聚体的主要特征。虽然用途与强聚(A)位点和弱聚(A)位点的关联不大,但DSE在强聚(A)位点附近更为明显。然后,我们使用包含六聚体和DSE的区域作为ERPIN程序识别PolyA位点的训练集,获得了69-85%的预测特异度,灵敏度为56%。全基因组和大型EST序列数据库的可获得性现在允许大规模观察多腺苷酸化位点。Poly(A)信号两侧的两个富含U的序列都有助于“真”位点的定义。然而,下游的富U序列也可能起到增强作用。基于这些信息,与以前最好的算法相比,Poly(A)位点预测的准确性得到了适度但持续的提高。
Differential polyadenylation is a widespread mechanism in higher eukaryotes producing mRNAs with different 3' ends in different contexts. This involves several alternative polyadenylation sites in the 3' UTR, each with its specific strength. Here, we analyze the vicinity of human polyadenylation signals in search of patterns that would help discriminate strong and weak polyadenylation sites, or true sites from randomly occurring signals. We used human genomic sequences to retrieve the region downstream of polyadenylation signals, usually absent from cDNA or mRNA databases. Analyzing 4956 EST-validated polyadenylation sites and their -300/+300 nt flanking regions, we clearly visualized the upstream (USE) and downstream (DSE) sequence elements, both characterized by U-rich (not GU-rich) segments. The presence of a USE and a DSE is the main feature distinguishing true polyadenylation sites from randomly occurring A(A/U)UAAA hexamers. While USEs are indifferently associated with strong and weak poly(A) sites, DSEs are more conspicuous near strong poly(A) sites. We then used the region encompassing the hexamer and DSE as a training set for poly(A) site identification by the ERPIN program and achieved a prediction specificity of 69 to 85% for a sensitivity of 56%. The availability of complete genomes and large EST sequence databases now permit large-scale observation of polyadenylation sites. Both U-rich sequences flanking both sides of poly(A) signals contribute to the definition of "true" sites. However, the downstream U-rich sequences may also play an enhancing role. Based on this information, poly(A) site prediction accuracy was moderately but consistently improved compared to the best previously available algorithm.
DOI: 10.1101/gr.190501
发表时间: 2001-09-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Beaudoing, E;Gautheret, D
通讯作者: Gautheret, D
DOI: 10.1093/nar/23.14.2614
发表时间: 1995-07-25
影响因子: 14.9
作者:
CHEN, F;MACDONALD, CC;WILUSZ, J
通讯作者: WILUSZ, J
DOI: 10.1093/nar/28.1.193
发表时间: 2000-01-01
影响因子: 14.9
作者:
Pesole, G;Liuni, S;Saccone, C
通讯作者: Saccone, C
DOI: 10.1006/jmbi.1997.0951
发表时间: 1997-04-25
影响因子: 5.6
作者:
Burge, C;Karlin, S
通讯作者: Karlin, S
DOI: 10.1038/ng780
发表时间: 2001-12-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Davuluri, RV;Grosse, I;Zhang, MQ
通讯作者: Zhang, MQ