Binding site graphs: a new graph theoretical framework for prediction of transcription factor binding sites.

Binding site graphs: a new graph theoretical framework for prediction of transcription factor binding sites.
复制标题

DOI:
10.1371/journal.pcbi.0030090
复制
发表时间:
2007-05
影响因子:
4.3
通讯作者:
Shakhnovich BE
Shakhnovich BE
中科院分区:
生物学2区
文献类型:
--
作者:
Reddy TE;DeLisi C;Shakhnovich BE

文献摘要

参考文献

相似文献

转录因子的核苷酸结合特异性的计算预测仍然是一个基本的和在很大程度上尚未解决的问题。确定结合位置是研究基因调控的先决条件,基因调控是控制表型多样性的主要机制。此外,从高通量数据源中准确确定结合特异性对于充分发挥系统生物学的潜力是必要的。不幸的是,最近进行的独立评估显示,来自最广泛使用的算法的预测有一半以上是错误的。我们引入了一个图论框架来描述局部序列相似性,即启动子序列中核苷酸之间的成对距离,并假设密集连接的子图指示转录因子结合位点。使用成熟的采样算法以及简单的聚类和评分方案,我们确定了几组密切相关的核苷酸,并测试了这些核苷酸的已知TF结合活性。使用一个独立的基准,我们发现我们的算法预测酵母结合基序比目前可用的技术要好得多,而且不需要人工管理。重要的是,我们将酵母中的假阳性预测数量减少到30%以下。我们还开发了一个框架来评估我们的主题预测的统计意义。我们表明,我们的方法对输入启动子的选择是稳健的,因此可以用于从噪声实验数据预测结合位置的背景下。我们应用我们的方法来利用基因组规模的芯片实验数据来识别结合位点。这些实验的结果可在http://cagt10.bu.edu/BSG.上公开获得。这里开发的图形框架在结合大量计算和实验测量的预测时可能会很有用。最后,我们讨论了如何使用我们的算法来提高对转录因子结合特异性的计算预测的敏感性。计算生物学中的一个历史难题是识别共调控基因启动子中的转录因子结合位点(TFBS)。随着对转录调控研究的日益重视,这个问题也与最近高通量和系统生物学实验的新结果具有独特的相关性。尽管在这一领域进行了广泛的研究,但最近对之前发表的技术的评估显示,还有很大的改进空间。在本文中,我们介绍了一种全新的识别TFBS的方法。首先,我们首先将启动子中的核苷酸表示为无向加权图。给定结合位点图(BSG)的这种表示,我们使用相对简单的图聚类技术来识别功能TFBS。我们表明,BSG预测在使用标准化评估基准的几乎所有性能衡量标准中都显著优于所有先前评估的方法。我们还发现,这种方法在选择输入启动子时比传统的Gibbs抽样更稳健,因此在有噪声的实验条件下更有可能表现得很好。最后,BSG非常擅长预测决定核苷酸的特异性。使用BSG预测,我们能够证实最近关于E-box转录因子CBF1和PHO4结合特异性的实验结果,并预测了TYE7的新的特异性决定核苷酸。
Computational prediction of nucleotide binding specificity for transcription factors remains a fundamental and largely unsolved problem. Determination of binding positions is a prerequisite for research in gene regulation, a major mechanism controlling phenotypic diversity. Furthermore, an accurate determination of binding specificities from high-throughput data sources is necessary to realize the full potential of systems biology. Unfortunately, recently performed independent evaluation showed that more than half the predictions from most widely used algorithms are false. We introduce a graph-theoretical framework to describe local sequence similarity as the pair-wise distances between nucleotides in promoter sequences, and hypothesize that densely connected subgraphs are indicative of transcription factor binding sites. Using a well-established sampling algorithm coupled with simple clustering and scoring schemes, we identify sets of closely related nucleotides and test those for known TF binding activity. Using an independent benchmark, we find our algorithm predicts yeast binding motifs considerably better than currently available techniques and without manual curation. Importantly, we reduce the number of false positive predictions in yeast to less than 30%. We also develop a framework to evaluate the statistical significance of our motif predictions. We show that our approach is robust to the choice of input promoters, and thus can be used in the context of predicting binding positions from noisy experimental data. We apply our method to identify binding sites using data from genome scale ChIP–chip experiments. Results from these experiments are publicly available at http://cagt10.bu.edu/BSG. The graphical framework developed here may be useful when combining predictions from numerous computational and experimental measures. Finally, we discuss how our algorithm can be used to improve the sensitivity of computational predictions of transcription factor binding specificities. A historically difficult problem in computational biology is the identification of transcription factor binding sites (TFBS) in the promoters of co-regulated genes. With increasing emphasis on research in transcriptional regulation, this problem is also uniquely relevant to emerging results from recent experiments in high-throughput and systems biology. Despite extensive research in the area, recent evaluations of previously published techniques show much room for improvement. In this paper, we introduce a fundamentally new approach to the identification of TFBS. First, we start by representing nucleotides in promoters as an undirected, weighted graph. Given this representation of a binding site graph (BSG), we employ relatively simple graph clustering techniques to identify functional TFBS. We show that BSG predictions significantly outperform all previously evaluated methods in nearly every performance measure using a standardized assessment benchmark. We also find that this approach is more robust than traditional Gibbs sampling to selection of input promoters, and thus more likely to perform well under noisy experimental conditions. Finally, BSGs are very good at predicting specificity determining nucleotides. Using BSG predictions, we were able to confirm recent experimental results on binding specificity of E-box TFs CBF1 and PHO4 and predict novel specificity determining nucleotides for TYE7.
DOI: 10.1093/nar/gkh169
发表时间: 2004-01-01
影响因子: 14.9
作者:
Frith, MC;Hansen, U;Weng, ZP
通讯作者: Weng, ZP
DOI: 10.1093/bioinformatics/btl243
发表时间: 2006-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Fratkin, Eugene;Naughton, Brian T.;Batzoglou, Serafim
通讯作者: Batzoglou, Serafim
DOI: 10.1126/science.1131007
发表时间: 2007-01-12
期刊: SCIENCE
影响因子: 56.9
作者:
Maerkl, Sebastian J.;Quake, Stephen R.
通讯作者: Quake, Stephen R.
DOI: 10.1186/1471-2105-6-84
发表时间: 2005-04-04
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Friberg, M;von Rohr, P;Gonnet, G
通讯作者: Gonnet, G
DOI: 10.1073/pnas.180265397
发表时间: 2000-08-29
影响因子: 11.1
作者:
Bussemaker, HJ;Li, H;Siggia, ED
通讯作者: Siggia, ED