Identification of transcription factor binding sites in the human genome sequence

Identification of transcription factor binding sites in the human genome sequence
复制标题

DOI:
10.1007/s00335-002-2175-6
复制
发表时间:
2002-09-01
期刊:
影响因子:
2.5
通讯作者:
Hannenhalli, S
Hannenhalli, S
中科院分区:
生物学4区
文献类型:
--
作者:
Levy, S;Hannenhalli, S

文献摘要

被引文献

相似文献

转录因子结合位点(TFBS)的鉴定是确定调控基因组转录的DNA信号的重要初始步骤。我们测试了应用于人类基因组序列的三种不同的TFBS识别计算方法的性能,根据它们恢复从TRANSFAC数据库中通过实验确定的和唯一映射的TFBS的位置的能力来判断。这些鉴定方法都试图通过比对描述结合位点的位置权重矩阵来过滤TFBS的数量,并采用(I)接受位点的P值阈值,(Ii)相邻位点的过度表示度量,或(Iii)与小鼠基因组的保守性和P值阈值的应用。结果表明,将人-鼠保守区的TFBS识别与非保守的非编码区TFBS识别相结合,可获得最好的TFBS识别效果。此外,我们发现,在481个实验定位的位点中,只有一半可以在小鼠保守的序列区域找到,但结合位点识别方法在保守区域的预测能力高达三倍。
The identification of transcription factor binding sites (TFBS) is an important initial step in determining the DNA signals that regulate transcription of the genome. We tested the performance of three distinct computational methods for the identification of TFBS applied to the human genome sequence, as judged by their ability to recover the location of experimentally determined, and uniquely mapped, TFBS taken from the TRANSFAC database. These identification methods all attempt to filter the quantity of TFBS identified by aligning positional weight matrices that describe the binding site and employ either (i) a P-value threshold for accepting a site, (ii) an over-representation measure of neighboring sites, or (iii) conservation with the mouse genome and application of P-value thresholds. The results show that the best recognition of TFBS is achieved by combining the identification of TFBS in regions of human-mouse conservation and also by applying a high stringency P-value to the TFBS identified in non-coding regions that are not conserved. Additionally, we find that only half of the 481 experimentally mapped sites can be found in sequence regions conserved with mouse, but the predictive power of the binding site identification method is up to threefold higher in the conserved regions.