Computational detection and location of transcription start sites in mammalian genomic DNA

Computational detection and location of transcription start sites in mammalian genomic DNA
复制标题

DOI:
10.1101/gr.216102
复制
发表时间:
2002-03-01
期刊:
影响因子:
7
通讯作者:
Hubbard, TJP
Hubbard, TJP
中科院分区:
生物学1区
文献类型:
--
作者:
Down, TA;Hubbard, TJP

文献摘要

被引文献

相似文献

转录,即从DNA基因组的片段中产生RNA拷贝的过程,由启动子区域指导。这些定义了转录起始位点,以及启动子有活性的细胞条件。至少在更复杂的物种中,基因具有几个不同的转录起始位点似乎是常见的,它们可能在不同的条件下具有活性。真核启动子是复杂且相当分散的结构,已证明难以在sillco中检测。我们证明了一种新的混合机器学习方法能够为>50%的人类转录起始位点构建有用的启动子模型。我们估计特异性为> 70%,并表现出良好的定位精度。基于我们学习的模型的结构,我们得出结论,类似于众所周知的TATA盒的信号,连同C-G富集的侧翼区域,是最重要的基于序列的信号,标记在一大类典型启动子的转录起始位点。
Transcription, the process whereby RNA copies are made from sections of the DNA genome, is directed by promoter regions. These define the transcription start site, and also the set of cellular conditions under which the promoter is active. At least in more complex species, it appears to be common for genes to have several different transcription start sites, which may be active under different conditions. Eukaryotic promoters are complex and fairly diffuse structures, which have proven hard to detect in sillco. We show that a novel hybrid machine-learning method is able to build useful models of promoters for >50% of human transcription start sites. We estimate specificity to be >70%, and demonstrate good positional accuracy. Based on the structure of our learned models, we conclude that a signal resembling the well known TATA box, together with flanking regions of C-G enrichment, are the most important sequence-based signals marking sites of transcriptional initiation at a large class of typical promoters.