Moment-based prediction of DNA-binding proteins

Moment-based prediction of DNA-binding proteins
复制标题

DOI:
10.1016/j.jmb.2004.05.058
复制
发表时间:
2004-07-30
影响因子:
5.6
通讯作者:
Sarai, A
Sarai, A
中科院分区:
生物学2区
文献类型:
--
作者:
Ahmad, S;Sarai, A

文献摘要

被引文献

相似文献

计算了62个具有代表性的已知结构的DNA结合蛋白的78个氨基酸序列的净电荷、电偶极矩和四极矩张量。结果发现,在这些链中的电荷分布的时刻的大小显着不同的非约束性的控制数据集。净电荷、净偶极矩和四极矩可以区分结合和非结合蛋白质,单变量预测的准确率分别为82.6%、77.4%和73.7%,无需交叉验证。使用混合预测与信息的电荷和两个时刻,最好的预测是85.6%,没有交叉验证和83.9%的交叉验证的数据集。使用这些简单描述符获得的这种预测准确度水平与使用包括许多描述符的更复杂模型获得的结果竞争。碳α原子上原子电荷的粗粒化并没有显著降低预测精度。这一结果表明,我们可以使用来自同源建模的C-α坐标来预测DNA结合蛋白。这种方法的速度和准确性,结合同源性为基础的结构预测方法,应提高全基因组识别的DNA结合蛋白。(C)2004 Elsevier Ltd.保留所有权利。
Net charge, electric dipole moment and quadrupole moment tensors were calculated for 78 amino acid sequences from 62 representative DNA-binding proteins with known structures. It was found that the magnitudes of the moments of electric charge distribution in these chains differ significantly from those of a non-binding control data set. Net charge, net dipole moment and quadrupole moment could each distinguish binding and non-binding proteins with 82.6%, 77.4% and 73.7% accuracy by single-variable predictors without cross-validation. Using hybrid predictors with information of charge and both moments, the best predictions were 85.6% without cross-validation and 83.9% for the cross-validated data sets. This level of prediction accuracy obtained with these simple descriptors competes with the results obtained using more complex models including many descriptors. The coarse graining of atomic charges onto C-alpha atoms did not reduce the prediction accuracy significantly. This result suggests that we can use C-alpha coordinates derived from homology modeling to predict DNA-binding proteins. The speed and accuracy of this method, in combination with homology-based methods of structure prediction, should enhance genome-wide recognition of DNA-binding proteins. (C) 2004 Elsevier Ltd. All rights reserved.