Statistical analysis of over-represented words in human promoter sequences

Statistical analysis of over-represented words in human promoter sequences
复制标题

DOI:
10.1093/nar/gkh246
复制
发表时间:
2004-02-01
影响因子:
14.9
通讯作者:
Landsman, D
Landsman, D
中科院分区:
生物学2区
文献类型:
--
作者:
Mariño-Ramírez, L;Spouge, JL;Landsman, D

文献摘要

被引文献

相似文献

了解转录起始位点(TSS)的精确位置,可以促进基因近端启动子区调控序列元件的鉴定和表征。利用5700多种不同的人类全长cdna中已知的tss,本研究从人类基因组中提取了4737个不同的推定启动子区域(PPRs)。相对于相应的TSS,每个PPR由-2000到+1000 bp的核苷酸组成。由于许多调控区域包含少于10个核苷酸的短而高度保守的字符串,我们在ppr中计算了8个字母的单词,使用z分数和其他相关统计来评估它们的过度和不足代表性。真核转录因子数据库TRANSFAC中描述了几个过度代表的八个字母单词的已知生物学功能;然而,许多人没有。除了使用与z分数相关的标准正态近似计算p值外,我们还使用了两个额外的统计控制来评估过度表示单词的重要性。这些控制对于用z分数评估过度和未充分代表的单词具有重要意义。
The identification and characterization of regulatory sequence elements in the proximal promoter region of a gene can be facilitated by knowing the precise location of the transcriptional start site (TSS). Using known TSSs from over 5700 different human full-length cDNAs, this study extracted a set of 4737 distinct putative promoter regions (PPRs) from the human genome. Each PPR consisted of nucleotides from -2000 to +1000 bp, relative to the corresponding TSS. Since many regulatory regions contain short, highly conserved strings of less than 10 nucleotides, we counted eight-letter words within the PPRs, using z-scores and other related statistics to evaluate their over- and under-representation. Several over-represented eight-letter words have known biological functions described in the eukaryotic transcription factor database TRANSFAC; however, many did not. Besides calculating a P-value with the standard normal approximation associated with z-scores, we used two extra statistical controls to evaluate the significance of over-represented words. These controls have important implications for evaluating over- and under-represented words with z-scores.