CART classification of human 5′ UTR sequences

CART classification of human 5′ UTR sequences
复制标题

DOI:
10.1101/gr.gr-1460r
复制
发表时间:
2000-11-01
期刊:
影响因子:
7
通讯作者:
Zhang, MQ
Zhang, MQ
中科院分区:
生物学1区
文献类型:
--
作者:
Davuluri, RV;Suzuki, Y;Zhang, MQ

文献摘要

被引文献

相似文献

使用最先进的实验和计算技术仔细制备了2312个全长人类5 '-非翻译区(UTR)的非冗余数据库。对该数据进行全面的计算分析以表征5' UTR特征。分类和回归树(CART)分析被用来将数据分为三个不同的类。I类由被认为翻译不良的mRNA组成,其具有充满潜在抑制特征的长5'UTR。II类由以生长依赖性方式调节的末端寡嘧啶段(TOP)mRNA组成,III类由具有可帮助有效翻译的有利5' UTR特征的mRNA组成。我们发现的最准确的树有92.5%的分类准确率,通过交叉验证估计。分类模型包括TOP的存在、二级结构、5' UTR长度和上游AUG(uAUG)的存在作为最相关的变量。目前对5'UTR的分类和表征为更好地理解人类mRNA的翻译调控提供了宝贵的信息。此外,该数据库和分类可以帮助人们建立更好的计算模型,用于预测5 '端外显子和将5' UTR与编码区分离。
A nonredundant database of 2312 full-length human 5'-untranslated regions (UTRs) was carefully prepared using state-of-the-art experimental and computational technologies. A comprehensive computational analysis of this data was conducted for characterizing the 5' UTR Features. Classification and regression tree (CART) analysis was used to classify the data into three distinct classes. Class I consists of mRNAs that are believed to be poorly translated with long 5' UTRs filled with potential inhibitory features. Class II consists of terminal oligopyrimidine tract (TOP) mRNAs that are regulated in a growth-dependent manner, and class III consists of mRNAs with Favorable 5' UTR features that may help efficient translation. The most accurate tree we found has 92.5% classification accuracy as estimated by cross validation. The classification model included the presence of TOP, a secondary structure, 5' UTR length, and the presence of upstream AUGs (uAUGs) as the most relevant variables. The present classification and characterization of the 5' UTRs provide precious information for better understanding the translational regulation of human mRNAs. Furthermore, this database and classification can help people build better computational models for predicting the 5'-terminal exon and separating the 5' UTR from the coding region.