Searching for statistically significant regulatory modules

Searching for statistically significant regulatory modules
复制标题

DOI:
10.1093/bioinformatics/btg1054
复制
发表时间:
2003-09-01
期刊:
影响因子:
5.8
通讯作者:
Noble, William Stafford
Noble, William Stafford
中科院分区:
生物学3区
文献类型:
--
作者:
Bailey, Timothy L.;Noble, William Stafford

文献摘要

被引文献

相似文献

动机:控制基因表达的调控机制是复杂的,常常需要多种同时发生的DNA -蛋白质相互作用。一个基因的转录速率可能取决于在基因附近的DNA上结合的一组转录因子的存在与否。在基因组DNA中定位转录因子结合位点是困难的,因为单个位点很小,而且往往偶然频繁出现。真正的结合位点可以通过它们成簇出现的倾向来识别,有时被称为调控模块。 结果:我们描述了一种用于检测基因组DNA中调控模块出现情况的算法。该算法称为MCAST,它将一个DNA数据库和一组已知协同作用的结合位点基序作为输入。MCAST使用一种基于基序的隐马尔可夫模型,该模型具有几个新颖的特征。该模型纳入了基序特异性的p值,从而能够直接比较不同宽度和特异性的基序的得分。p值评分还允许MCAST只接受显著性低于用户指定阈值的基序出现情况,同时仍然给p值更低的基序出现情况赋予更好的分数。MCAST可以搜索长的DNA序列,对调控模块内基序之间的长度分布进行建模,但忽略模块之间的长度分布。该算法产生一个按E值排序的预测调控模块列表。我们使用模拟数据以及来自果蝇和人类的真实数据集对该算法进行了验证。
Motivation: The regulatory machinery controlling gene expression is complex, frequently requiring multiple, simultaneous DNA-protein interactions. The rate at which a gene is transcribed may depend upon the presence or absence of a collection of transcription factors bound to the DNA near the gene. Locating transcription factor binding sites in genomic DNA is difficult because the individual sites are small and tend to occur frequently by chance. True binding sites may be identified by their tendency to occur in clusters, sometimes known as regulatory modules.Results: We describe an algorithm for detecting occurrences of regulatory modules in genomic DNA. The algorithm, called MCAST, takes as input a DNA database and a collection of binding site motifs that are known to operate in concert. MCAST uses a motif-based hidden Markov model with several novel features. The model incorporates motif-specific p-values, thereby allowing scores from motifs of different widths and specificities to be compared directly. The p-value scoring also allows MCAST to only accept motif occurrences with significance below a user-specified threshold, while still assigning better scores to motif occurrences with lower p-values. MCAST can search long DNA sequences, modeling length distributions between motifs within a regulatory module, but ignoring length distributions between modules. The algorithm produces a list of predicted regulatory modules, ranked by E-value. We validate the algorithm using simulated data as well as real data sets from fruitfly and human.