Sequence information for the splicing of human Pre-mRNA identified by support vector machine classification

Sequence information for the splicing of human Pre-mRNA identified by support vector machine classification
复制标题

DOI:
10.1101/gr.1679003
复制
发表时间:
2003-12-01
期刊:
影响因子:
7
通讯作者:
Chasin, LA
Chasin, LA
中科院分区:
生物学1区
文献类型:
--
作者:
Zhang, XHF;Heller, KA;Chasin, LA

文献摘要

被引文献

相似文献

脊椎动物mrna前转录物包含许多类似于剪接位点的序列,但这些更多的假剪接位点通常被细胞剪接机制完全忽略。即使在外显子定义的层面上,由这种假剪接位点定义的伪外显子数量也比真实外显子多一个数量级。我们使用支持向量机来发现序列信息,这些信息可以用来区分真实外显子和伪外显子。该机器学习工具定义了潜在分支点,扩展的多嘧啶束,以及在组成剪接外显子上游50 nt的区域内富含c和富含tg的基序。在外显子下游80 nt的区域也发现了富含c的序列,以及g三联体基序。此外,剪接供体一致序列中三个碱基的组合比一致值在区分真实剪接位点和伪剪接位点方面更有效;双向碱基组合是区分3'剪接位点的最佳选择。这些数据还表明,两个或多个这些元件之间的相互作用可能有助于外显子识别,并提供候选序列作为内含子剪接增强子进行评估。
Vertebrate pre-mRNA transcripts contain many sequences that resemble splice sites on the basis of agreement to the consensus, yet these more numerous false splice sites are usually completely ignored by the cellular splicing machinery. Even at the level of exon definition, pseudo exons defined by such false splices sites outnumber real exons by an order of magnitude. We used a support vector machine to discover sequence information that could be used to distinguish real exons from pseudo exons. This machine learning tool led to the definition of potential branch points, an extended polypyrimidine tract, and C-rich and TG-rich motifs in a region limited to 50 nt upstream of constitutively spliced exons. C-rich sequences were also found in a region extending to 80 nt downstream of exons, along with G-triplet motifs. In addition, it was shown that combinations of three bases within the splice donor consensus sequence were more effective than consensus values in distinguishing real from pseudo splice sites; two-way base combinations were optimal for distinguishing 3' splice sites. These data also suggest that interactions between two or more of these elements may contribute to exon recognition, and provide candidate sequences for assessment as intronic splicing enhancers.