NeBcon: protein contact map prediction using neural network training coupled with naiive Bayes classifiers

NeBcon: protein contact map prediction using neural network training coupled with naiive Bayes classifiers
复制标题

DOI:
10.1093/bioinformatics/btx164
复制
发表时间:
2017-08-01
期刊:
影响因子:
5.8
通讯作者:
Zhang, Yang
Zhang, Yang
中科院分区:
生物学3区
文献类型:
--
作者:
He, Baoji;Mortuza, S. M.;Zhang, Yang

文献摘要

被引文献

相似文献

动机:最近的CASP实验见证了在基于联合进化的接触预测的帮助下折叠大尺寸非单体蛋白质的令人兴奋的进展。然而,由于接触预测方法需要大量的序列同源序列,而大多数非单体蛋白靶标无法获得这些序列同源序列,因此成功的预测是坊间传闻。为了提高从头算蛋白质结构预测的成功率,开发高效的方法来为不同类型的蛋白质靶标生成平衡可靠的接触图是至关重要的。结果:我们开发了一种新的流水线NeBcon,它使用朴素贝叶斯分类器(NBC)定理结合了从协同进化和机器学习方法建立的8种最先进的接触方法。然后通过神经网络学习将NBC模型的后验概率与固有结构特征一起训练,以得到最终的接触图预测。NeBcon在98个非冗余蛋白质上进行了测试,它将基于最佳协同进化的元服务器预测器的准确率提高了22%;对于数据库中缺乏序列和结构同源性的硬目标,改进的幅度增加到45%。详细的数据分析表明,改进的主要贡献是来自协同进化和机器学习预测的补充信息的NBC优化组合。神经网络训练还有助于改善NBC后验概率和内在结构特征的耦合,这被发现对没有足够数量的同源序列来获得可靠的共同进化图谱的蛋白质特别重要。
Motivation: Recent CASP experiments have witnessed exciting progress on folding large-size non-humongous proteins with the assistance of co-evolution based contact predictions. The success is however anecdotal due to the requirement of the contact prediction methods for the high volume of sequence homologs that are not available to most of the non-humongous protein targets. Development of efficient methods that can generate balanced and reliable contact maps for different type of protein targets is essential to enhance the success rate of the ab initio protein structure prediction.Results: We developed a new pipeline, NeBcon, which uses the naiive Bayes classifier (NBC) theorem to combine eight state of the art contact methods that are built from co-evolution and machine learning approaches. The posterior probabilities of the NBC model are then trained with intrinsic structural features through neural network learning for the final contact map prediction. NeBcon was tested on 98 non-redundant proteins, which improves the accuracy of the best co-evolution based meta-server predictor by 22%; the magnitude of the improvement increases to 45% for the hard targets that lack sequence and structural homologs in the databases. Detailed data analysis showed that the major contribution to the improvement is due to the optimized NBC combination of the complementary information from both co-evolution and machine learning predictions. The neural network training also helps to improve the coupling of the NBC posterior probability and the intrinsic structural features, which were found particularly important for the proteins that do not have sufficient number of homologous sequences to derive reliable co-evolution profiles.