DeepQA: improving the estimation of single protein model quality with deep belief networks.

DeepQA: improving the estimation of single protein model quality with deep belief networks.
复制标题

DOI:
10.1186/s12859-016-1405-y
复制
发表时间:
2016-12-05
期刊:
影响因子:
3
通讯作者:
Cheng J
Cheng J
中科院分区:
生物学4区
文献类型:
--
作者:
Cao R;Bhattacharya D;Hou J;Cheng J

文献摘要

参考文献

被引文献

相似文献

长期以来,蛋白质质量评价(QA)一直被认为是蛋白质三级结构预测的主要挑战之一。特别是,估计单个蛋白质模型的质量,这对于从由大多数低质量模型组成的大型模型池中选择几个好的模型很重要,仍然是一个很大程度上未解决的问题。本文提出了一种基于深度信念网络的单模型质量评估方法DeepQA,该方法利用能量、物理化学特征和结构信息等多个特征从不同角度描述模型的质量。深度信念网络在几个大型数据集上进行训练,这些数据集包括来自蛋白质结构预测关键评估(CASP)实验的模型、几个公开可用的数据集和我们内部从头开始方法生成的模型。我们的实验表明,深度信念网络在蛋白质模型质量评估问题上比支持向量机和神经网络具有更好的性能,并且我们的方法DeepQA在CASP11数据集上达到了最先进的性能。在从从头开始建模方法生成的大量低质量模型中选择良好的离群模型方面,它也优于两种成熟的方法。DeepQA是一个有用的深度学习工具,用于蛋白质单模型质量评估和蛋白质结构预测。DeepQA的源代码、可执行文件、文档和训练/测试数据集对非商业用户免费提供,网址为http://cactus.rnet.missouri.edu/DeepQA/。本文的在线版本(doi:10.1186/s12859-016-1405-y)包含补充材料,授权用户可以使用。
Protein quality assessment (QA) useful for ranking and selecting protein models has long been viewed as one of the major challenges for protein tertiary structure prediction. Especially, estimating the quality of a single protein model, which is important for selecting a few good models out of a large model pool consisting of mostly low-quality models, is still a largely unsolved problem. We introduce a novel single-model quality assessment method DeepQA based on deep belief network that utilizes a number of selected features describing the quality of a model from different perspectives, such as energy, physio-chemical characteristics, and structural information. The deep belief network is trained on several large datasets consisting of models from the Critical Assessment of Protein Structure Prediction (CASP) experiments, several publicly available datasets, and models generated by our in-house ab initio method. Our experiments demonstrate that deep belief network has better performance compared to Support Vector Machines and Neural Networks on the protein model quality assessment problem, and our method DeepQA achieves the state-of-the-art performance on CASP11 dataset. It also outperformed two well-established methods in selecting good outlier models from a large set of models of mostly low quality generated by ab initio modeling methods. DeepQA is a useful deep learning tool for protein single model quality assessment and protein structure prediction. The source code, executable, document and training/test datasets of DeepQA for Linux is freely available to non-commercial users at http://cactus.rnet.missouri.edu/DeepQA/. The online version of this article (doi:10.1186/s12859-016-1405-y) contains supplementary material, which is available to authorized users.
DOI: 10.1038/srep25687
发表时间: 2016-05-10
期刊: Scientific reports
影响因子: 4.6
作者:
Li J;Cheng J
通讯作者: Cheng J
DOI: 10.1186/s12859-015-0775-x
发表时间: 2015-10-23
期刊: BMC bioinformatics
影响因子: 3
作者:
Li J;Cao R;Cheng J
通讯作者: Cheng J
DOI: 10.1093/bioinformatics/btq662
发表时间: 2011-02-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Benkert P;Biasini M;Schwede T
通讯作者: Schwede T
DOI: 10.1186/1472-6807-14-13
发表时间: 2014-04-15
影响因子: --
作者:
Cao R;Wang Z;Cheng J
通讯作者: Cheng J
DOI: 10.1038/srep23990
发表时间: 2016-04-04
期刊: Scientific reports
影响因子: 4.6
作者:
Cao R;Cheng J
通讯作者: Cheng J