Accurate prediction of protein disordered regions by mining protein structure data

Accurate prediction of protein disordered regions by mining protein structure data
复制标题

DOI:
10.1007/s10618-005-0001-y
复制
发表时间:
2005-11-01
影响因子:
4.8
通讯作者:
Baldi, P
Baldi, P
中科院分区:
计算机科学3区
文献类型:
--
作者:
Cheng, JL;Sweredoski, MJ;Baldi, P

文献摘要

被引文献

相似文献

蛋白质中的内含子无序区是相对常见的,对于我们理解分子识别和组装以及蛋白质结构和功能非常重要。从算法的角度来看,标记大的无序区域对于从头计算蛋白质结构预测方法也很重要。在这里,我们首先从蛋白质数据库中提取一个精选的、非冗余的蛋白质无序区域数据集,并计算这些区域的长度和位置的相关统计数据。然后,我们开发了一个从头预测的无序区域称为DISpro,它使用的形式的配置文件,预测的二级结构和相对溶剂的可及性,和集成的一维递归神经网络的进化信息。DISpro使用精选的数据集进行训练和交叉验证。实验结果表明,DISpro的准确率为92.8%,假阳性率为5%。DISpro是SCRATCH蛋白质数据挖掘工具套件的成员,可通过http://www.igb.uci.edu/servers/psss.html获得。
Intrinsically disordered regions in proteins are relatively frequent and important for our understanding of molecular recognition and assembly, and protein structure and function. From an algorithmic standpoint, flagging large disordered regions is also important for ab initio protein structure prediction methods. Here we first extract a curated, non-redundant, data set of protein disordered regions from the Protein Data Bank and compute relevant statistics on the length and location of these regions. We then develop an ab initio predictor of disordered regions called DISpro which uses evolutionary information in the form of profiles, predicted secondary structure and relative solvent accessibility, and ensembles of 1D-recursive neural networks. DISpro is trained and cross validated using the curated data set. The experimental results show that DISpro achieves an accuracy of 92.8% with a false positive rate of 5%. DISpro is a member of the SCRATCH suite of protein data mining tools available through http://www.igb.uci.edu/servers/psss.html.