Identification of DNA-binding proteins using support vector machines and evolutionary profiles.

Identification of DNA-binding proteins using support vector machines and evolutionary profiles.
复制标题

DOI:
10.1186/1471-2105-8-463
复制
发表时间:
2007-11-27
期刊:
影响因子:
3
通讯作者:
Raghava GP
Raghava GP
中科院分区:
生物学4区
文献类型:
--
作者:
Kumar M;Gromiha MM;Raghava GP

文献摘要

参考文献

被引文献

相似文献

DNA结合蛋白的鉴定是基因组注释领域的主要挑战之一,因为这些蛋白在基因调控中起着至关重要的作用。在本文中,我们开发了各种支持向量机模块预测DNA结合域和蛋白质。所有模型都在非冗余蛋白质的多个数据集上进行了训练和测试。在由1153个DNA结合蛋白和等量非DNA结合蛋白组成的DNAaset上建立了SVM模型,在氨基酸和二肽组分上分别获得了72.42%和71.59%的最大准确率。当以PSSM图谱形式的进化信息代替氨基酸组成作为输入时,SVM模型的性能从72.42%提高到74.22%。此外,支持向量机模型已被开发的DNAset,其中包括146个DNA结合和250个非结合链/结构域,并取得了最大的准确性79.80%和86.62%,使用氨基酸组成和PSSM的配置文件。本研究中开发的SVM模型在盲数据集上的性能优于现有方法。一个高度准确的方法已被开发用于预测DNA结合蛋白质使用SVM和PSSM配置文件。这是第一个研究中,进化信息的PSSM配置文件的形式已被成功地用于预测DNA结合蛋白。一个网络服务器DNAbinder已被开发用于识别DNA结合蛋白质和结构域从查询氨基酸序列。
Identification of DNA-binding proteins is one of the major challenges in the field of genome annotation, as these proteins play a crucial role in gene-regulation. In this paper, we developed various SVM modules for predicting DNA-binding domains and proteins. All models were trained and tested on multiple datasets of non-redundant proteins. SVM models have been developed on DNAaset, which consists of 1153 DNA-binding and equal number of non DNA-binding proteins, and achieved the maximum accuracy of 72.42% and 71.59% using amino acid and dipeptide compositions, respectively. The performance of SVM model improved from 72.42% to 74.22%, when evolutionary information in form of PSSM profiles was used as input instead of amino acid composition. In addition, SVM models have been developed on DNAset, which consists of 146 DNA-binding and 250 non-binding chains/domains, and achieved the maximum accuracy of 79.80% and 86.62% using amino acid composition and PSSM profiles. The SVM models developed in this study perform better than existing methods on a blind dataset. A highly accurate method has been developed for predicting DNA-binding proteins using SVM and PSSM profiles. This is the first study in which evolutionary information in form of PSSM profiles has been used successfully for predicting DNA-binding proteins. A web-server DNAbinder has been developed for identifying DNA-binding proteins and domains from query amino acid sequences .
DOI: 10.1093/nar/gki949
发表时间: 2005
影响因子: 14.9
作者:
Bhardwaj N;Langlois RE;Zhao G;Lu H
通讯作者: Lu H
DOI: 10.1093/bioinformatics/bth322
发表时间: 2004-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Kaur, H;Raghava, GPS
通讯作者: Raghava, GPS
DOI: 10.1110/ps.0241703
发表时间: 2003-05-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
Kaur, H;Raghava, GPS
通讯作者: Raghava, GPS
DOI: 10.1110/ps.0228903
发表时间: 2003-03-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
Kaur, H;Raghava, GPS
通讯作者: Raghava, GPS
DOI: 10.1016/j.jmb.2004.05.058
发表时间: 2004-07-30
影响因子: 5.6
作者:
Ahmad, S;Sarai, A
通讯作者: Sarai, A