Improving the accuracy of protein secondary structure prediction using structural alignment

Improving the accuracy of protein secondary structure prediction using structural alignment
复制标题

DOI:
10.1186/1471-2105-7-301
复制
发表时间:
2006-06-14
期刊:
影响因子:
3
通讯作者:
Wishart, David S.
Wishart, David S.
中科院分区:
生物学4区
文献类型:
--
作者:
Montgomerie, Scott;Sundararaj, Shan;Wishart, David S.

文献摘要

被引文献

相似文献

背景:在过去30年中,蛋白质二级结构预测的准确性稳步提高。现在,许多二级结构预测方法通常达到约75%的精度(Q3)。我们认为,通过将结构(相对于序列)数据库比较作为预测过程的一部分,可以进一步提高这种准确性。实际上,鉴于蛋白质数据库的较大尺寸(> 35,000个序列),具有结构同源物的新鉴定序列的概率实际上很高。分析:我们开发了一种方法,该方法可以执行基于结构的序列比对作为作为一部分的一部分。二级结构预测过程。通过将已知同源物(序列ID> 25%)的结构映射到查询蛋白序列上,可以预测该查询蛋白二次结构的至少一部分。通过将这种结构比对方法与常规(基于序列的)二级结构方法相结合,然后将其与“陪审团”系统结合起来,以产生共识结果,可以实现非常高的预测准确性。使用EVA的1644蛋白的序列唯一测试集,这种新方法的平均Q3得分为81.3%。广泛的测试表明,这比当前可用的任何其他方法都要好4-5%。使用非序列唯一的测试集的评估(蛋白质体注释或结构基因组学的典型代表)表明,这种新方法可以达到Q3得分接近88%。结论:通过同时使用序列和结构数据库,并利用最新技术机器学习可以常规地预测蛋白质二级结构,其精度远高于80%。可以在http://wishart.biology.ualberta.ca/proteus上访问一个名为Proteus的程序和Web服务器,该程序和Web服务器可执行这些二级结构预测。对于高吞吐量或批处理序列分析,可以在本地下载并运行Proteus程序,数据库(和服务器)。
Background: The accuracy of protein secondary structure prediction has steadily improved over the past 30 years. Now many secondary structure prediction methods routinely achieve an accuracy (Q3) of about 75%. We believe this accuracy could be further improved by including structure (as opposed to sequence) database comparisons as part of the prediction process. Indeed, given the large size of the Protein Data Bank (> 35,000 sequences), the probability of a newly identified sequence having a structural homologue is actually quite high.Results: We have developed a method that performs structure-based sequence alignments as part of the secondary structure prediction process. By mapping the structure of a known homologue (sequence ID > 25%) onto the query protein's sequence, it is possible to predict at least a portion of that query protein's secondary structure. By integrating this structural alignment approach with conventional (sequence-based) secondary structure methods and then combining it with a "jury-of-experts" system to generate a consensus result, it is possible to attain very high prediction accuracy. Using a sequence-unique test set of 1644 proteins from EVA, this new method achieves an average Q3 score of 81.3%. Extensive testing indicates this is approximately 4 - 5% better than any other method currently available. Assessments using non sequence-unique test sets (typical of those used in proteome annotation or structural genomics) indicate that this new method can achieve a Q3 score approaching 88%.Conclusion: By using both sequence and structure databases and by exploiting the latest techniques in machine learning it is possible to routinely predict protein secondary structure with an accuracy well above 80%. A program and web server, called PROTEUS, that performs these secondary structure predictions is accessible at http://wishart.biology.ualberta.ca/proteus. For high throughput or batch sequence analyses, the PROTEUS programs, databases (and server) can be downloaded and run locally.