Identifying target sites for cooperatively binding factors

Identifying target sites for cooperatively binding factors
复制标题

DOI:
10.1093/bioinformatics/17.7.608
复制
发表时间:
2001-07-01
期刊:
影响因子:
5.8
通讯作者:
Stormo, GD
Stormo, GD
中科院分区:
生物学3区
文献类型:
--
作者:
GuhaThakurta, D;Stormo, GD

文献摘要

被引文献

相似文献

动机:真核生物中的转录激活通常需要多种转录因子的组合相互作用。虽然存在几种方法用于鉴定DNA序列中的单个蛋白质结合位点模式,但用于发现协同作用因子的结合位点模式的方法很少。在这里,我们提出了一个算法,合作绑定(合作绑定),发现DNA靶位点的协同作用的转录因子。该方法利用吉布斯采样策略来模拟两个转录因子之间的协同性,并定义结合位点的位置权重矩阵。从训练集和整个基因组的序列被考虑在内,以区别于常见的模式在基因组中,并产生模式,这是显着的,只有在training set.Results:我们已经测试了共绑定半合成和真实的数据集,以显示它可以有效地识别DNA靶位点的模式,协同结合转录因子。在结合位点模式较弱且无法通过其他可用方法识别的情况下,Co-Bind凭借对因素之间的协同性建模,可以有效地识别这些位点。虽然开发了蛋白质-DNA相互作用的模型,但Co-Bind的范围可以扩展到其他大分子中的组合、序列特异性相互作用。
Motivation: Transcriptional activation in eukaryotic organisms normally requires combinatorial interactions of multiple transcription factors. Though several methods exist for identification of individual protein binding site patterns in DNA sequences, there are few methods for discovery of binding site patterns for cooperatively acting factors. Here we present an algorithm, Co-Bind (for COperative BINDing), for discovering DNA target sites for cooperatively acting transcription factors. The method utilizes a Gibbs sampling strategy to model the cooperativity between two transcription factors and defines position weight matrices for the binding sites. Sequences from both the training set and the entire genome are taken into account, in order to discriminate against commonly occurring patterns in the genome, and produce patterns which are significant only in the training set.Results: We have tested Co-Bind on semi-synthetic and real data sets to show it can efficiently identify DNA target site patterns for cooperatively binding transcription factors. In cases where binding site patterns are weak and cannot be identified by other available methods, Co-Bind, by virtue of modeling the cooperativity between factors, can identify those sites efficiently. Though developed to model protein-DNA interactions, the scope of Co-Bind may be extended to combinatorial, sequence specific, interactions in other macromolecules.