Prediction of RNA-binding proteins from primary sequence by a support vector machine approach

Prediction of RNA-binding proteins from primary sequence by a support vector machine approach
复制标题

DOI:
10.1261/rna.5890304
复制
发表时间:
2004-03-01
期刊:
RNA
影响因子:
4.5
通讯作者:
Chen, YZ
Chen, YZ
中科院分区:
生物学3区
文献类型:
--
作者:
Han, LY;Cai, CZ;Chen, YZ

文献摘要

被引文献

相似文献

阐明蛋白质与不同分子的相互作用对于理解细胞过程具有重要意义。已经开发了用于预测蛋白质-蛋白质相互作用的计算方法。但对蛋白质-RNA相互作用的预测却关注不够,而蛋白质-RNA相互作用在调控基因表达和某些RNA介导的酶促过程中起着核心作用。这项工作探索了使用机器学习方法,支持向量机(SVM),直接从它们的一级序列预测RNA结合蛋白。基于已知的RNA结合蛋白和非RNA结合蛋白的知识,训练SVM系统以识别RNA结合蛋白。共使用4011个RNA结合蛋白和9781个非RNA结合蛋白来训练和测试SVM分类系统,并且使用447个RNA结合蛋白和4881个非RNA结合蛋白的独立集合来评估分类准确性。使用该独立评估集的测试结果显示,rRNA、mRNA和tRNA结合蛋白的预测准确度分别为94.1%、79.3%和94.1%,非rRNA、非mRNA和非tRNA结合蛋白的预测准确度分别为98.7%、96.5%和99.9%。对一小类只有60个可用序列的snRNA结合蛋白进一步测试了SVM分类系统。对于snRNA结合蛋白和非snRNA结合蛋白的预测准确率分别为40.0%和99.9%,这表明需要足够数量的蛋白来训练SVM。在这项工作中训练的SVM分类系统被添加到我们基于Web的蛋白质功能分类软件SVMProt中,网址为http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi。我们的研究表明,支持向量机作为一个有用的工具,促进蛋白质-RNA相互作用的预测的潜力。
Elucidation of the interaction of proteins with different molecules is of significance in the understanding of cellular processes. Computational methods have been developed for the prediction of protein-protein interactions. But insufficient attention has been paid to the prediction of protein-RNA interactions, which play central roles in regulating gene expression and certain RNA-mediated enzymatic processes. This work explored the use of a machine learning method, support vector machines (SVM), for the prediction of RNA-binding proteins directly from their primary sequence. Based on the knowledge of known RNA-binding and non-RNA-binding proteins, an SVM system was trained to recognize RNA-binding proteins. A total of 4011 RNA-binding and 9781 non-RNA-binding proteins was used to train and test the SVM classification system, and an independent set of 447 RNA-binding and 4881 non-RNA-binding proteins was used to evaluate the classification accuracy. Testing results using this independent evaluation set show a prediction accuracy of 94.1%, 79.3%, and 94.1% for rRNA-, mRNA-, and tRNA-binding proteins, and 98.7%, 96.5%, and 99.9% for non-rRNA-, non-mRNA-, and non-tRNA-binding proteins, respectively. The SVM classification system was further tested on a small class of snRNA-binding proteins with only 60 available sequences. The prediction accuracy is 40.0% and 99.9% for snRNA-binding and non-snRNA-binding proteins, indicating a need for a sufficient number of proteins to train SVM. The SVM classification systems trained in this work were added to our Web-based protein functional classification software SVMProt, at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi. Our study suggests the potential of SVM as a useful tool for facilitating the prediction of protein-RNA interactions.