Connecting protein structure with predictions of regulatory sites

Connecting protein structure with predictions of regulatory sites
复制标题

DOI:
10.1073/pnas.0701356104
复制
发表时间:
2007-04-24
影响因子:
11.1
通讯作者:
Siggia, Eric D.
Siggia, Eric D.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Morozov, Alexandre V.;Siggia, Eric D.

文献摘要

被引文献

相似文献

微阵列实验提出的一个共同任务是从一组已知转录因子所调节的基因中推断出其结合位点的偏好,并确定该因子是单独作用还是在一个复合体中起作用。也可以提出相反的问题:给定一组结合位点,可以推断出调节因子或因子的复合体吗?通过使用相对简单的蛋白质- dna相互作用的同源模型以及快速扩展的蛋白质结构数据库,这两项任务都得到了极大的便利。对于出芽酵母,我们能够构建67个转录因子的可靠结构模型,并通过使用贝叶斯吉布斯采样算法和广泛的蛋白质定位数据集重新确定因子结合位点。对于49个与该数据集的先前分析(主要基于系统发育保守)相同的因素,我们发现先前预测的结合基序中有一半需要进行一些修改。我们还解决了从结合位点确定因子的逆向问题,通过分配正确的蛋白质折叠到先前研究中的49例中的25例。我们的方法很容易扩展到其他生物,包括高等真核生物。我们的研究强调了扩大当前结构基因组学项目的效用,这些项目详尽地采样折叠结构空间,以包括具有显著不同dna结合特异性的所有因素。
A common task posed by microarray experiments is to infer the binding site preferences for a known transcription factor from a collection of genes that it regulates and to ascertain whether the factor acts alone or in a complex. The converse problem can also be posed: Given a collection of binding sites, can the regulatory factor or complex of factors be inferred? Both tasks are substantially facilitated by using relatively simple homology models for protein-DNA interactions, as well as the rapidly expanding protein structure database. For budding yeast, we are able to construct reliable structural models for 67 transcription factors and with them redetermine factor binding sites by using a Bayesian Gibbs sampling algorithm and an extensive protein localization data set. For 49 factors in common with a prior analysis of this data set (based largely on phylogenetic conservation), we find that half of the previously predicted binding motifs are in need of some revision. We also solve the inverse problem of ascertaining the factors from the binding sites by assigning a correct protein fold to 25 of the 49 cases from a previous study. Our approach is easily extended to other organisms, including higher eukaryotes. Our study highlights the utility of enlarging current structural genomics projects that exhaustively sample fold structure space to include all factors with significantly different DNA-binding specificities.