Predicting RNA-binding sites of proteins using support vector machines and evolutionary information.

Predicting RNA-binding sites of proteins using support vector machines and evolutionary information.
复制标题

DOI:
10.1186/1471-2105-9-s12-s6
复制
发表时间:
2008-12-12
期刊:
影响因子:
3
通讯作者:
Hsu WL
Hsu WL
中科院分区:
生物学4区
文献类型:
--
作者:
Cheng CW;Su EC;Hwang JK;Sung TY;Hsu WL

文献摘要

被引文献

相似文献

RNA-蛋白质相互作用在蛋白质合成、基因表达、转录后调控和病毒感染性等生物学过程中发挥着重要作用。蛋白质中RNA结合位点的识别为生物学家提供了有价值的见解。然而,RNA-蛋白质相互作用的实验测定仍然费时费力。因此,预测蛋白质中RNA结合位点的计算方法变得非常可取。对RNA结合位点预测的广泛研究导致了几种方法的发展。然而,它们在为高特异性进行权衡时可能会产生低灵敏度。我们提出了一种方法,RNAProB,它结合了一种新的平滑的位置特定评分矩阵(PSSM)编码方案和支持向量机模型来预测蛋白质中的RNA结合位点。除了结合标准PSSM谱的进化信息外,提出的平滑PSSM编码方案还考虑了蛋白质中每个氨基酸与邻近残基的相关性和依赖性。实验结果表明,平滑PSSM编码显著提高了预测性能,尤其是灵敏度。使用五次交叉验证,我们的方法在总体准确度、特异度和马太相关系数方面分别比最先进的系统提高了4.90%~6.83%、0.88%~5.33%和0.10~0.23。最值得注意的是,与其他方法相比,RNAProB在基准数据集上显著提高了7.0%~26.9%的灵敏度。为了防止数据过拟合,引入了三向数据分裂过程来估计预测性能。此外,还对RNA结合蛋白的理化性质和氨基酸偏好进行了检测和分析。我们的结果表明,平滑的PSSM编码方案显著提高了蛋白质中RNA结合位点的预测性能。这也支持了我们的假设,即平滑的PSSM编码通过对周围残基的依赖进行建模,可以更好地解决区分相互作用和非相互作用残基的模糊性。该方法还可用于DNA结合位点预测、蛋白质相互作用、翻译后修饰位点预测等研究领域。
RNA-protein interaction plays an essential role in several biological processes, such as protein synthesis, gene expression, posttranscriptional regulation and viral infectivity. Identification of RNA-binding sites in proteins provides valuable insights for biologists. However, experimental determination of RNA-protein interaction remains time-consuming and labor-intensive. Thus, computational approaches for prediction of RNA-binding sites in proteins have become highly desirable. Extensive studies of RNA-binding site prediction have led to the development of several methods. However, they could yield low sensitivities in trade-off for high specificities. We propose a method, RNAProB, which incorporates a new smoothed position-specific scoring matrix (PSSM) encoding scheme with a support vector machine model to predict RNA-binding sites in proteins. Besides the incorporation of evolutionary information from standard PSSM profiles, the proposed smoothed PSSM encoding scheme also considers the correlation and dependency from the neighboring residues for each amino acid in a protein. Experimental results show that smoothed PSSM encoding significantly enhances the prediction performance, especially for sensitivity. Using five-fold cross-validation, our method performs better than the state-of-the-art systems by 4.90%~6.83%, 0.88%~5.33%, and 0.10~0.23 in terms of overall accuracy, specificity, and Matthew's correlation coefficient, respectively. Most notably, compared to other approaches, RNAProB significantly improves sensitivity by 7.0%~26.9% over the benchmark data sets. To prevent data over fitting, a three-way data split procedure is incorporated to estimate the prediction performance. Moreover, physicochemical properties and amino acid preferences of RNA-binding proteins are examined and analyzed. Our results demonstrate that smoothed PSSM encoding scheme significantly enhances the performance of RNA-binding site prediction in proteins. This also supports our assumption that smoothed PSSM encoding can better resolve the ambiguity of discriminating between interacting and non-interacting residues by modelling the dependency from surrounding residues. The proposed method can be used in other research areas, such as DNA-binding site prediction, protein-protein interaction, and prediction of posttranslational modification sites.