Tree-based position weight matrix approach to model transcription factor binding site profiles.

Tree-based position weight matrix approach to model transcription factor binding site profiles.
复制标题

DOI:
10.1371/journal.pone.0024210
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Davuluri RV
Davuluri RV
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Bi Y;Kim H;Gupta R;Davuluri RV

文献摘要

参考文献

被引文献

相似文献

大多数用于预测转录因子结合位点 (TFBS) 的基于位置权重矩阵 (PWM) 的生物信息学方法都假设序列基序中的每个核苷酸独立地对蛋白质和 DNA 序列之间的相互作用做出贡献,通常会产生较高的假阳性预测。最近的 ChIP-Seq 方法越来越多地提供 TF 富集图谱,这有助于研究 TFBS 的依赖结构和准确预测。我们开发了一种新颖的基于树的 PWM (TPWM) 方法来准确模拟 TF 与其结合位点之间的相互作用。整个树形结构 PWM 可以被视为不同条件 PWM 的混合。我们提出了一种称为 TPD(基于 TPWM 的判别方法)的判别方法,利用预先存在的 PWM 从 ChIP-Seq 数据构建 TPWM。为了实现正负数据集之间的最大判别力,根据马太相关系数(MCC)确定截止值。所得到的 TPWM 在广泛的合成数据集上进行准确性评估。然后,我们将 TPWM 判别方法应用于几个真实的 ChIP-Seq 数据集,以细化 TRANSFAC 数据库中存储的当前 TFBS 模型。对模拟和真实 ChIP-Seq 数据的实验表明,所提出的从现有 PWM 出发的方法在检测 TFBS 方面始终比现有工具具有更好的性能。准确性的提高是对主题的完整依赖结构进行建模以及更好地预测真阳性率的结果。这些发现可能有助于更好地理解 TF-DNA 相互作用的机制。
Most of the position weight matrix (PWM) based bioinformatics methods developed to predict transcription factor binding sites (TFBS) assume each nucleotide in the sequence motif contributes independently to the interaction between protein and DNA sequence, usually producing high false positive predictions. The increasing availability of TF enrichment profiles from recent ChIP-Seq methodology facilitates the investigation of dependent structure and accurate prediction of TFBSs. We develop a novel Tree-based PWM (TPWM) approach to accurately model the interaction between TF and its binding site. The whole tree-structured PWM could be considered as a mixture of different conditional-PWMs. We propose a discriminative approach, called TPD (TPWM based Discriminative Approach), to construct the TPWM from the ChIP-Seq data with a pre-existing PWM. To achieve the maximum discriminative power between the positive and negative datasets, the cutoff value is determined based on the Matthew Correlation Coefficient (MCC). The resulting TPWMs are evaluated with respect to accuracy on extensive synthetic datasets. We then apply our TPWM discriminative approach on several real ChIP-Seq datasets to refine the current TFBS models stored in the TRANSFAC database. Experiments on both the simulated and real ChIP-Seq data show that the proposed method starting from existing PWM has consistently better performance than existing tools in detecting the TFBSs. The improved accuracy is the result of modelling the complete dependent structure of the motifs and better prediction of true positive rate. The findings could lead to better understanding of the mechanisms of TF-DNA interactions.
DOI: 10.1093/bfgp/elp014
发表时间: 2009-07-01
期刊: Briefings in Functional Genomics & Proteomics
影响因子: --
作者:
Narlikar, Leelavati;Ovcharenko, Ivan
通讯作者: Ovcharenko, Ivan
DOI: 10.1093/nar/30.5.1255
发表时间: 2002-03-01
影响因子: 14.9
作者:
Bulyk, ML;Johnson, PLF;Church, GM
通讯作者: Church, GM
DOI: 10.1093/nar/gkn488
发表时间: 2008-09
影响因子: 14.9
作者:
Jothi, Raja;Cuddapah, Suresh;Barski, Artem;Cui, Kairong;Zhao, Keji
通讯作者: Zhao, Keji
DOI: 10.1186/gb-2009-10-11-r131
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Essien K;Vigneau S;Apreleva S;Singh LN;Bartolomei MS;Hannenhalli S
通讯作者: Hannenhalli S
DOI: 10.1101/gr.089086.108
发表时间: 2009-06-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Bruce, Alexander W.;Lopez-Contreras, Andres J.;Vetrie, David
通讯作者: Vetrie, David