Sequence and structural features of carbohydrate binding in proteins and assessment of predictability using a neural network

Sequence and structural features of carbohydrate binding in proteins and assessment of predictability using a neural network
复制标题

DOI:
10.1186/1472-6807-7-1
复制
发表时间:
2007-01-03
影响因子:
--
通讯作者:
Ahmad, Shandar
Ahmad, Shandar
中科院分区:
生物4区
文献类型:
--
作者:
Malik, Adeel;Ahmad, Shandar

文献摘要

被引文献

相似文献

背景:蛋白质-碳水化合物相互作用在许多生物学过程中至关重要,并对药物靶向和基因表达产生影响。蛋白质-碳水化合物相互作用的性质可以在单个残基水平上通过分析结合区中与非结合区相比的局部序列和结构环境来研究,这为此类分析提供了内在对照。最终目的是从序列和结构预测结合位点,需要编制结合区域的总体统计数据。基于序列的结合位点预测在我们早期的工作中已经成功地应用于DNA结合蛋白。我们的目标是将类似的分析应用于碳水化合物结合蛋白。然而,由于参与这种相互作用的蛋白质区域相对较小,因此方法和结果显着不同。蛋白质-碳水化合物复合物的比较也与其他蛋白质-配体complex.Results:我们已经编制了统计数据的氨基酸组合物的结合与非结合区域的一般以及在每个不同的二级结构构象。结合倾向的20个残基类型和它们的结构特征,如溶剂的可及性,堆积密度和二级结构已被计算,以评估其倾向于碳水化合物相互作用。最后,氨基酸序列的进化概况已被用于使用神经网络预测结合位点。使用来自单个序列的信息训练另一组神经网络,并比较来自进化谱和单个序列的预测性能。最好的基于神经网络的预测可以实现87%的预测灵敏度和23%的特异性,所有碳水化合物结合位点,使用进化信息。单一序列对同一数据集的敏感性为68%,特异性为55%。有限的半乳糖结合数据集的灵敏度和特异性分别为63%和79%的进化信息和62%和68%的灵敏度和特异性的单一序列。碳水化合物结合位点的倾向和其他序列和结构特征也进行了比较与我们类似的广泛的研究DNA结合蛋白质,也与蛋白质配体complex.Conclusion:碳水化合物通常表现出偏好结合芳香族残基,最突出的色氨酸。结合位点的较高暴露表面积表明疏水相互作用的作用。神经网络给出了适度的成功预测,预计在未来更多的蛋白质-碳水化合物复合物的结构变得可用时,这将得到改善。
Background: Protein-Carbohydrate interactions are crucial in many biological processes with implications to drug targeting and gene expression. Nature of protein-carbohydrate interactions may be studied at individual residue level by analyzing local sequence and structure environments in binding regions in comparison to non-binding regions, which provide an inherent control for such analyses. With an ultimate aim of predicting binding sites from sequence and structure, overall statistics of binding regions needs to be compiled. Sequence-based predictions of binding sites have been successfully applied to DNA-binding proteins in our earlier works. We aim to apply similar analysis to carbohydrate binding proteins. However, due to a relatively much smaller region of proteins taking part in such interactions, the methodology and results are significantly different. A comparison of protein-carbohydrate complexes has also been made with other protein-ligand complexes.Results: We have compiled statistics of amino acid compositions in binding versus non-binding regions-general as well as in each different secondary structure conformation. Binding propensities of each of the 20 residue types and their structure features such as solvent accessibility, packing density and secondary structure have been calculated to assess their predisposition to carbohydrate interactions. Finally, evolutionary profiles of amino acid sequences have been used to predict binding sites using a neural network. Another set of neural networks was trained using information from single sequences and the prediction performance from the evolutionary profiles and single sequences were compared. Best of the neural network based prediction could achieve an 87% sensitivity of prediction at 23% specificity for all carbohydrate-binding sites, using evolutionary information. Single sequences gave 68% sensitivity and 55% specificity for the same data set. Sensitivity and specificity for a limited galactose binding data set were obtained as 63% and 79% respectively for evolutionary information and 62% and 68% sensitivity and specificity for single sequences. Propensity and other sequence and structural features of carbohydrate binding sites have also been compared with our similar extensive studies on DNA-binding proteins and also with protein-ligand complexes.Conclusion: Carbohydrates typically show a preference to bind aromatic residues and most prominently tryptophan. Higher exposed surface area of binding sites indicates a role of hydrophobic interactions. Neural networks give a moderate success of prediction, which is expected to improve when structures of more protein-carbohydrate complexes become available in future.