High-throughput SELEX-SAGE method for quantitative modeling of transcription-factor binding sites

High-throughput SELEX-SAGE method for quantitative modeling of transcription-factor binding sites
复制标题

DOI:
10.1038/nbt718
复制
发表时间:
2002-08-01
影响因子:
46.9
通讯作者:
Bucher, P
Bucher, P
中科院分区:
工程技术1区
文献类型:
--
作者:
Roulet, E;Busso, S;Bucher, P

文献摘要

被引文献

相似文献

确定基因组中所有转录因子结合位点的位置和相对强度的能力对于全面了解基因调控和生物技术应用中有效的启动子工程都很重要。在这里,我们提出了一种生物信息驱动的实验方法来准确定义转录因子的 DNA 结合序列特异性。广义谱(1)用作结合位点的预测定量模型,其参数是使用标准隐马尔可夫模型训练算法从体外选择的配体估计的(2,3)。计算机模拟表明,需要数千个低到中等亲和力序列才能生成所需精度的概况。为了产生如此规模的数据,我们将高通量基因组学方法应用于此处解决的生化问题。将指数富集配体系统进化 (SELEX)(4) 和基因表达系列分析 (SAGE)(5) 方案相结合的方法与基于 Phred 质量评分 (6) 的自动质量控制序列提取程序相结合。这使得能够对包含 10,000 多个 CTF/NFI 转录因子潜在 DNA 配体的数据库进行测序。由此产生的结合位点模型以前所未有的高精度定义了该蛋白质的序列特异性,从而使识别基因组 DNA 中以前未知的调控序列成为可能。所选位点的协方差分析揭示了不同核苷酸位置的非独立碱基偏好,从而提供了对结合机制的深入了解。
The ability to determine the location and relative strength of all transcription-factor binding sites in a genome is important both for a comprehensive understanding of gene regulation and for effective promoter engineering in biotechnological applications. Here we present a bioinformatically driven experimental method to accurately define the DNA-binding sequence specificity of transcription factors. A generalized profile(1) was used as a predictive quantitative model for binding sites, and its parameters were estimated from in vitro-selected ligands using standard hidden Markov model training algorithms(2,3). Computer simulations showed that several thousand low- to medium-affinity sequences are required to generate a profile of desired accuracy. To produce data on this scale, we applied high-throughput genomics methods to the biochemical problem addressed here. A method combining systematic evolution of ligands by exponential enrichment (SELEX)(4) and serial analysis of gene expression (SAGE)(5) protocols was coupled to an automated quality-controlled sequence extraction procedure based on Phred quality scores(6). This allowed the sequencing of a database of more than 10,000 potential DNA ligands for the CTF/NFI transcription factor. The resulting binding-site model defines the sequence specificity of this protein with a high degree of accuracy not achieved earlier and thereby makes it possible to identify previously unknown regulatory sequences in genomic DNA. A covariance analysis of the selected sites revealed non-independent base preferences at different nucleotide positions, providing insight into the binding mechanism.