Prediction of mono- and di-nucleotide-specific DNA-binding sites in proteins using neural networks.

Prediction of mono- and di-nucleotide-specific DNA-binding sites in proteins using neural networks.
复制标题

DOI:
10.1186/1472-6807-9-30
复制
发表时间:
2009-05-13
影响因子:
--
通讯作者:
Ahmad S
Ahmad S
中科院分区:
生物4区
文献类型:
--
作者:
Andrabi M;Mizuguchi K;Sarai A;Ahmad S

文献摘要

参考文献

被引文献

相似文献

蛋白质对 DNA 的识别是生命系统中最重要的过程之一。因此,了解一般的识别过程,特别是识别蛋白质和DNA中的相互识别位点具有重要意义。蛋白质中 DNA 结合位点的序列和结构依赖性导致了用于预测的成功机器学习方法的开发。然而,所有现有的机器学习方法都预测 DNA 结合位点,无论其目标序列如何,因此,它们都无助于识别特定的蛋白质-DNA 接触。在这项工作中,我们提出了根据蛋白质残基环境与 DNA 中单核苷酸或二核苷酸步骤之间的接触来预测特定 DNA 结合位点的问题。这项工作的目的是将蛋白质序列或结构特征作为输入,并预测每个氨基酸残基是否在四种可能的单核苷酸之一或 10 个独特的二核苷酸步骤之一识别的位置处与 DNA 结合。接触预测是在不同的分辨率级别上进行的。就 DNA 的侧链、主链和大沟或小沟原子而言。观察到特定接触的残基偏好存在显着差异,这与其他特征相结合,导致了有希望的预测水平。一般来说,基于 PSSM 的预测,在二级结构和溶剂可及性的支持下,可达到约 70-80% 的良好预测性,通过 ROC 图的曲线下面积 (AUC) 来测量。大沟和小沟接触预测的突出之处在于其序列或 PSSM 的可预测性较差,通过添加二级结构和溶剂可及性信息可以非常有效地(> 20 个百分点)补偿,揭示了局部蛋白质结构在大/小沟 DNA 识别中的主导作用。对结果进行详细分析后,开发了一个使用 PSSM 预测单核苷酸和二核苷酸步骤接触的网络服务器,并在 或 上提供。仅使用序列和进化信息就可以高精度地预测大多数残基-核苷酸接触。然而,主要和次要凹槽接触很大程度上取决于局部结构。总体而言,这项研究使我们离预测蛋白质和 DNA 序列中相互识别位点的最终目标又近了一步。
DNA recognition by proteins is one of the most important processes in living systems. Therefore, understanding the recognition process in general, and identifying mutual recognition sites in proteins and DNA in particular, carries great significance. The sequence and structural dependence of DNA-binding sites in proteins has led to the development of successful machine learning methods for their prediction. However, all existing machine learning methods predict DNA-binding sites, irrespective of their target sequence and hence, none of them is helpful in identifying specific protein-DNA contacts. In this work, we formulate the problem of predicting specific DNA-binding sites in terms of contacts between the residue environments of proteins and the identity of a mononucleotide or a dinucleotide step in DNA. The aim of this work is to take a protein sequence or structural features as inputs and predict for each amino acid residue if it binds to DNA at locations identified by one of the four possible mononucleotides or one of the 10 unique dinucleotide steps. Contact predictions are made at various levels of resolution viz. in terms of side chain, backbone and major or minor groove atoms of DNA. Significant differences in residue preferences for specific contacts are observed, which combined with other features, lead to promising levels of prediction. In general, PSSM-based predictions, supported by secondary structure and solvent accessibility, achieve a good predictability of ~70–80%, measured by the area under the curve (AUC) of ROC graphs. The major and minor groove contact predictions stood out in terms of their poor predictability from sequences or PSSM, which was very strongly (>20 percentage points) compensated by the addition of secondary structure and solvent accessibility information, revealing a predominant role of local protein structure in the major/minor groove DNA-recognition. Following a detailed analysis of results, a web server to predict mononucleotide and dinucleotide-step contacts using PSSM was developed and made available at or . Most residue-nucleotide contacts can be predicted with high accuracy using only sequence and evolutionary information. Major and minor groove contacts, however, depend profoundly on the local structure. Overall, this study takes us a step closer to the ultimate goal of predicting mutual recognition sites in protein and DNA sequences.
DOI: 10.1093/nar/gki949
发表时间: 2005
影响因子: 14.9
作者:
Bhardwaj N;Langlois RE;Zhao G;Lu H
通讯作者: Lu H
DOI: 10.1016/j.cell.2005.10.042
发表时间: 2006-01-13
期刊: CELL
影响因子: 64.5
作者:
Hallikas, O;Palin, K;Taipale, J
通讯作者: Taipale, J
DOI: 10.1016/j.jmb.2004.05.058
发表时间: 2004-07-30
影响因子: 5.6
作者:
Ahmad, S;Sarai, A
通讯作者: Sarai, A
DOI: 10.1038/nsb0994-638
发表时间: 1994-09-01
期刊: NATURE STRUCTURAL BIOLOGY
影响因子: --
作者:
KIM, JL;BURLEY, SK
通讯作者: BURLEY, SK
DOI: 10.1093/nar/gkj131
发表时间: 2006-01-01
影响因子: 14.9
作者:
Kummerfeld SK;Teichmann SA
通讯作者: Teichmann SA