Novel approach for selecting the best predictor for identifying the binding sites in DNA binding proteins.

Novel approach for selecting the best predictor for identifying the binding sites in DNA binding proteins.
复制标题

DOI:
10.1093/nar/gkt544
复制
发表时间:
2013-09
影响因子:
14.9
通讯作者:
Gromiha MM
Gromiha MM
中科院分区:
生物学2区
文献类型:
--
作者:
Nagarajan R;Ahmad S;Gromiha MM

文献摘要

参考文献

被引文献

相似文献

蛋白质-DNA复合体通过氨基酸与DNA的相互作用在许多细胞过程中发挥重要作用。已经发展了几种计算方法来利用序列和/或结构信息来预测DNA结合蛋白中的相互作用残基。这些方法显示了不同程度的准确度,这可能取决于训练中使用的数据集的选择、为开发预测模型而选择的特征集、模型获取对预测有用的信息的能力或这些因素的组合。在许多情况下,不同的方法可能会产生类似的结果,而在其他情况下,预测者可能会返回相互矛盾的预测。在这种情况下,对适用于所研究系统的预测性能的先验估计将有助于生物学家选择最佳方法来设计他们的实验。在这项工作中,我们基于各种生物学上的相关考虑,构建了无偏的、严格的和多样化的DNA结合蛋白数据集:(I)7个结构类别,(Ii)86个折叠,(Iii)106个超家族,(Iv)194个家族,(V)15个结合基序,(Vi)单链/双链DNA,(Vii)DNA构象(A,B,Z等),(Viii)三个功能和(Ix)无序区。这些数据集被剔除为非冗余的,序列同一性分别为25%和40%,并用于评估在线服务或独立程序可用的11种不同方法的性能。我们观察到,每个数据集的最佳执行方法都显示出对为其基准选择的数据集的显著偏见。我们的分析揭示了重要的数据集特征,这些特征可以用来估计这些特定于上下文的偏差,从而建议用于给定问题的最佳方法。我们已经开发了一个Web服务器,它根据需要考虑这些功能,并显示调查人员应该使用的最佳方法。该Web服务器可在http://www.biotech.iitm.ac.in/DNA-protein/.免费获得在此基础上,根据算法的复杂度对算法进行了分组,并对算法的性能进行了分析。本工作所获得的信息可以有效地用于选择最佳的实验设计方法。
Protein–DNA complexes play vital roles in many cellular processes by the interactions of amino acids with DNA. Several computational methods have been developed for predicting the interacting residues in DNA-binding proteins using sequence and/or structural information. These methods showed different levels of accuracies, which may depend on the choice of data sets used in training, the feature sets selected for developing a predictive model, the ability of the models to capture information useful for prediction or a combination of these factors. In many cases, different methods are likely to produce similar results, whereas in others, the predictors may return contradictory predictions. In this situation, a priori estimates of prediction performance applicable to the system being investigated would be helpful for biologists to choose the best method for designing their experiments. In this work, we have constructed unbiased, stringent and diverse data sets for DNA-binding proteins based on various biologically relevant considerations: (i) seven structural classes, (ii) 86 folds, (iii) 106 superfamilies, (iv) 194 families, (v) 15 binding motifs, (vi) single/double-stranded DNA, (vii) DNA conformation (A, B, Z, etc.), (viii) three functions and (ix) disordered regions. These data sets were culled as non-redundant with sequence identities of 25 and 40% and used to evaluate the performance of 11 different methods in which online services or standalone programs are available. We observed that the best performing methods for each of the data sets showed significant biases toward the data sets selected for their benchmark. Our analysis revealed important data set features, which could be used to estimate these context-specific biases and hence suggest the best method to be used for a given problem. We have developed a web server, which considers these features on demand and displays the best method that the investigator should use. The web server is freely available at http://www.biotech.iitm.ac.in/DNA-protein/. Further, we have grouped the methods based on their complexity and analyzed the performance. The information gained in this work could be effectively used to select the best method for designing experiments.
DOI: 10.1002/jcc.10009
发表时间: 2002-01-15
影响因子: 3
作者:
Jayaram, B;McConnell, K;Beveridge, DL
通讯作者: Beveridge, DL
DOI: 10.1016/j.febslet.2007.01.086
发表时间: 2007-03-06
期刊: FEBS LETTERS
影响因子: 3.5
作者:
Bhardwaj, Nitin;Lu, Hui
通讯作者: Lu, Hui
DOI: 10.1038/329263a0
发表时间: 1987-09-17
期刊: NATURE
影响因子: 64.8
作者:
HOGAN, ME;AUSTIN, RH
通讯作者: AUSTIN, RH
DOI: 10.1016/s0006-3495(92)81649-1
发表时间: 1992-09-01
影响因子: 3.4
作者:
BERMAN, HM;OLSON, WK;SCHNEIDER, B
通讯作者: SCHNEIDER, B
DOI: 10.1093/bioinformatics/btl672
发表时间: 2007-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hwang, Seungwoo;Gou, Zhenkun;Kuznetsov, Igor B.
通讯作者: Kuznetsov, Igor B.