High Resolution Models of Transcription Factor-DNA Affinities Improve In Vitro and In Vivo Binding Predictions

High Resolution Models of Transcription Factor-DNA Affinities Improve In Vitro and In Vivo Binding Predictions
复制标题

DOI:
10.1371/journal.pcbi.1000916
复制
发表时间:
2010-09-01
影响因子:
4.3
通讯作者:
Leslie, Christina
Leslie, Christina
中科院分区:
生物学2区
文献类型:
--
作者:
Agius, Phaedra;Arvey, Aaron;Leslie, Christina

文献摘要

被引文献

相似文献

准确地模拟转录因子(TF)的DNA序列偏好,并使用这些模型来预测TF的体内基因组结合位点,是破译调控密码的关键。TF结合位点基序的可用性和准确性有限,通常表示为位置特异性评分矩阵(PSSM),这可能会匹配大量的网站,并产生一个不可靠的目标基因列表,这些努力都受到了挫折。最近,蛋白结合微阵列(PBM)实验已经成为一个新的来源,高分辨率的数据在体外TF结合特异性。PBM数据已经通过估计PSSM或通过对探针强度的秩统计进行了分析,使得个体序列模式被分配富集分数(E分数)。这种表示是信息丰富但不实用的,因为每个TF被分配了数千个评分序列模式的列表。同时,来自ChIP-seq实验的高分辨率体内TF占用数据也越来越多。我们已经开发了一个灵活的判别框架,用于从高分辨率的体外和体内数据中学习TF结合偏好。我们首先在PBM数据上训练支持向量回归(SVR)模型,以学习从探针序列到结合强度的映射。我们使用了一种新的k-mer为基础的字符串内核称为双错配内核来表示探针序列的相似性。SVR模型比E-评分更紧凑,比PSSM更有表达力,并且可以容易地用于扫描基因组区域以预测体内占用。使用酵母和小鼠TF的大数据集,我们发现我们的SVR模型可以更好地预测探针强度比E-score方法或PBM-derived PSSM。此外,通过使用SVRs对酵母、小鼠和人类基因组区域进行评分,我们能够更好地预测通过ChIP-chip和ChIP-seq实验测量的基因组占用率。最后,我们发现,通过直接在ChIP-seq数据上训练基于内核的模型,我们大大提高了体内占有率预测,并且通过比较TF的体外和体内模型,我们可以识别辅因子并消除直接和间接结合的歧义。
Accurately modeling the DNA sequence preferences of transcription factors (TFs), and using these models to predict in vivo genomic binding sites for TFs, are key pieces in deciphering the regulatory code. These efforts have been frustrated by the limited availability and accuracy of TF binding site motifs, usually represented as position-specific scoring matrices (PSSMs), which may match large numbers of sites and produce an unreliable list of target genes. Recently, protein binding microarray (PBM) experiments have emerged as a new source of high resolution data on in vitro TF binding specificities. PBM data has been analyzed either by estimating PSSMs or via rank statistics on probe intensities, so that individual sequence patterns are assigned enrichment scores (E-scores). This representation is informative but unwieldy because every TF is assigned a list of thousands of scored sequence patterns. Meanwhile, high-resolution in vivo TF occupancy data from ChIP-seq experiments is also increasingly available. We have developed a flexible discriminative framework for learning TF binding preferences from high resolution in vitro and in vivo data. We first trained support vector regression (SVR) models on PBM data to learn the mapping from probe sequences to binding intensities. We used a novel k-mer based string kernel called the di-mismatch kernel to represent probe sequence similarities. The SVR models are more compact than E-scores, more expressive than PSSMs, and can be readily used to scan genomics regions to predict in vivo occupancy. Using a large data set of yeast and mouse TFs, we found that our SVR models can better predict probe intensity than the E-score method or PBM-derived PSSMs. Moreover, by using SVRs to score yeast, mouse, and human genomic regions, we were better able to predict genomic occupancy as measured by ChIP-chip and ChIP-seq experiments. Finally, we found that by training kernel-based models directly on ChIP-seq data, we greatly improved in vivo occupancy prediction, and by comparing a TF's in vitro and in vivo models, we could identify cofactors and disambiguate direct and indirect binding.