Three-stage prediction of protein β-sheets by neural networks, alignments and graph algorithms

Three-stage prediction of protein β-sheets by neural networks, alignments and graph algorithms
复制标题

DOI:
10.1093/bioinformatics/bti1004
复制
发表时间:
2005-06-01
期刊:
影响因子:
5.8
通讯作者:
Baldi, P
Baldi, P
中科院分区:
生物学3区
文献类型:
--
作者:
Cheng, JL;Baldi, P

文献摘要

被引文献

相似文献

动机:蛋白质β折叠在蛋白质结构、功能、进化和生物工程中起着基础作用。然而,蛋白质β-折叠的准确预测和组装仍然具有挑战性,因为蛋白质β-折叠需要在线性距离的残基之间形成氢键。以前的方法用于预测β-片层的拓扑特征,如β-链比对,一般没有利用的全球协变和约束特性的β-片层architecture.Results:我们提出了一个模块化的方法来预测/组装蛋白质β-片层在一个链中的问题,通过整合本地和全球的约束,在三个步骤。第一步使用递归神经网络来预测配对概率的所有对链间β-残基的概况,二级结构和溶剂可及性信息。第二步将动态规划技术应用于这些概率,以获得结合假能和所有β链对之间的最佳对齐。最后,第三步使用图匹配算法来预测蛋白质的β-折叠结构,通过优化全局伪能量,同时实施强全局β-链配对约束。该方法使用交叉验证方法在一个大型的非同源数据集进行评估,并产生显着的改进,比以前的方法。
Motivation: Protein beta-sheets play a fundamental role in protein structure, function, evolution and bioengineering. Accurate prediction and assembly of protein beta-sheets, however, remains challenging because protein beta-sheets require formation of hydrogen bonds between linearly distant residues. Previous approaches for predicting beta-sheet topological features, such as beta-strand alignments, in general have not exploited the global covariation and constraints characteristic of beta-sheet architectures.Results: We propose a modular approach to the problem of predicting/assembling protein beta-sheets in a chain by integrating both local and global constraints in three steps. The first step uses recursive neural networks to predict pairing probabilities for all pairs of interstrand beta-residues from profile, secondary structure and solvent accessibility information. The second step applies dynamic programming techniques to these probabilities to derive binding pseudoenergies and optimal alignments between all pairs of beta-strands. Finally, the third step uses graph matching algorithms to predict the beta-sheet architecture of the protein by optimizing the global pseudoenergy while enforcing strong global beta-strand pairing constraints. The approach is evaluated using cross-validation methods on a large non-homologous dataset and yields significant improvements over previous methods.