DBD: a transcription factor prediction database.

DBD: a transcription factor prediction database.
复制标题

DOI:
10.1093/nar/gkj131
复制
发表时间:
2006-01-01
影响因子:
14.9
通讯作者:
Teichmann SA
Teichmann SA
中科院分区:
生物学2区
文献类型:
--
作者:
Kummerfeld SK;Teichmann SA

文献摘要

参考文献

被引文献

相似文献

基因表达的调控影响生物体中几乎所有的生物过程;序列特异性DNA结合转录因子对这种控制至关重要。对于大多数基因组,转录因子库仅部分已知。Hierarchy转录因子鉴定主要基于使用成对序列比较的基因组注释管道,其仅检测与已知基因相似的那些因子,或者基于将许多类型的蛋白质合并到“转录因子”类别中的功能分类方案。使用一种新的转录因子鉴定方法,DBD转录因子数据库填补了这一空白,为整个生命树的生物体提供全基因组转录因子预测。DBD背后的预测方法通过使用结构域的轮廓隐马尔可夫模型(Profile Hidden Markov Models,简称HRM)的同源性来识别序列特异性DNA结合转录因子。因此,它仅限于与那些HISTORY同源的因子。HIBRARY的集合来自两个现有的数据库(Pfam和SUPERFAMILY),并且仅限于专门检测特异性识别DNA序列的转录因子的模型。例如,它不包括基础转录因子或染色质相关蛋白。基于与实验验证的注释的比较,预测过程的准确度在95%和99%之间。在我们全基因组预测的转录因子中,有四分之一到一半代表了以前未表征的蛋白质。DBD()包括150个完全测序的基因组的预测转录因子库,它们的结构域分配和手工策划的DNA结合结构域障碍列表。用户可以通过基因组、结构域家族或序列标识符浏览、搜索或下载预测,基于结构域架构查看转录因子家族,并接收蛋白质序列的预测。
Regulation of gene expression influences almost all biological processes in an organism; sequence-specific DNA-binding transcription factors are critical to this control. For most genomes, the repertoire of transcription factors is only partially known. Hitherto transcription factor identification has been largely based on genome annotation pipelines that use pairwise sequence comparisons, which detect only those factors similar to known genes, or on functional classification schemes that amalgamate many types of proteins into the category of ‘transcription factor’. Using a novel transcription factor identification method, the DBD transcription factor database fills this void, providing genome-wide transcription factor predictions for organisms from across the tree of life. The prediction method behind DBD identifies sequence-specific DNA-binding transcription factors through homology using profile hidden Markov models (HMMs) of domains. Thus, it is limited to factors that are homologus to those HMMs. The collection of HMMs is taken from two existing databases (Pfam and SUPERFAMILY), and is limited to models that exclusively detect transcription factors that specifically recognize DNA sequences. It does not include basal transcription factors or chromatin-associated proteins, for instance. Based on comparison with experimentally verified annotation, the prediction procedure is between 95% and 99% accurate. Between one quarter and one-half of our genome-wide predicted transcription factors represent previously uncharacterized proteins. The DBD () consists of predicted transcription factor repertoires for 150 completely sequenced genomes, their domain assignments and the hand curated list of DNA-binding domain HMMs. Users can browse, search or download the predictions by genome, domain family or sequence identifier, view families of transcription factors based on domain architecture and receive predictions for a protein sequence.
DOI: 10.1093/nar/gkh117
发表时间: 2004-01-01
影响因子: 14.9
作者:
Madera, M;Vogel, C;Gough, J
通讯作者: Gough, J
DOI: 10.1093/nar/gki046
发表时间: 2005-01-01
影响因子: 14.9
作者:
Drysdale RA;Crosby MA;FlyBase Consortium
通讯作者: FlyBase Consortium
DOI: 10.1093/nar/gkh012
发表时间: 2004-01-01
影响因子: 14.9
作者:
Sandelin, A;Alkema, W;Lenhard, B
通讯作者: Lenhard, B
DOI: 10.1126/science.1075090
发表时间: 2002-10-25
期刊: SCIENCE
影响因子: 56.9
作者:
Lee, TI;Rinaldi, NJ;Young, RA
通讯作者: Young, RA
DOI: 10.1186/gb-2002-3-3-research0012
发表时间: 2002
期刊: Genome biology
影响因子: 12.3
作者:
Iyer LM;Koonin EV;Aravind L
通讯作者: Aravind L