Dinucleotide weight matrices for predicting transcription factor binding sites: generalizing the position weight matrix.

Dinucleotide weight matrices for predicting transcription factor binding sites: generalizing the position weight matrix.
复制标题

DOI:
10.1371/journal.pone.0009722
复制
发表时间:
2010-03-22
期刊:
影响因子:
3.7
通讯作者:
Siddharthan R
Siddharthan R
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Siddharthan R

文献摘要

参考文献

被引文献

相似文献

通过计算机识别转录因子结合位点(TFBS)是理解基因调控的关键。TFBS是表现出一些可变性的字符串模式,通常被建模为“位置权重矩阵”(PWMs)。虽然方便,PWM具有显着的局限性,特别是假设的独立性的结合基序内的位置;和基于PWM的预测通常不是非常具体的已知的功能位点。在酵母中结合位点的分析表明,二核苷酸的相关性不仅限于近邻,但可以扩展到相当大的差距。我描述了一个简单的PWM模型的推广,认为频率的二核苷酸,而不是个别核苷酸。与以前的努力不同,该方法考虑了扩展的结合区域内的所有二核苷酸,并且不试图先验地确定特定二核苷酸相关性的意义。我描述了如何使用“二核苷酸权重矩阵”(DWM)来预测结合位点,特别是处理的复杂性,它的条目不是独立的概率。基准测试表明,在许多因素下,PWM在预测已知目标的精度方面有了显着的提高。在大多数情况下,显著的进一步改善是通过将通常定义的“核心基序”在两侧延伸约10 bp来实现的。虽然该侧翼序列在核苷酸水平上没有显示出强基序,但二核苷酸模型的预测能力表明,DNA序列中蛋白质结合亲和力的“签名”延伸到核心蛋白质-DNA接触区域之外。虽然计算上比基于PWM的方法要求更高,速度更慢,但这种二核苷酸方法在概念上和实现上都很简单,可以作为未来改进的基础。
Identifying transcription factor binding sites (TFBS) in silico is key in understanding gene regulation. TFBS are string patterns that exhibit some variability, commonly modelled as “position weight matrices” (PWMs). Though convenient, the PWM has significant limitations, in particular the assumed independence of positions within the binding motif; and predictions based on PWMs are usually not very specific to known functional sites. Analysis here on binding sites in yeast suggests that correlation of dinucleotides is not limited to near-neighbours, but can extend over considerable gaps. I describe a straightforward generalization of the PWM model, that considers frequencies of dinucleotides instead of individual nucleotides. Unlike previous efforts, this method considers all dinucleotides within an extended binding region, and does not make an attempt to determine a priori the significance of particular dinucleotide correlations. I describe how to use a “dinucleotide weight matrix” (DWM) to predict binding sites, dealing in particular with the complication that its entries are not independent probabilities. Benchmarks show, for many factors, a dramatic improvement over PWMs in precision of predicting known targets. In most cases, significant further improvement arises by extending the commonly defined “core motifs” by about 10bp on either side. Though this flanking sequence shows no strong motif at the nucleotide level, the predictive power of the dinucleotide model suggests that the “signature” in DNA sequence of protein-binding affinity extends beyond the core protein-DNA contact region. While computationally more demanding and slower than PWM-based approaches, this dinucleotide method is straightforward, both conceptually and in implementation, and can serve as a basis for future improvements.
DOI: 10.1371/journal.pcbi.0010067
发表时间: 2005-12
影响因子: 4.3
作者:
Siddharthan, Rahul;Siggia, Eric D;van Nimwegen, Erik
通讯作者: van Nimwegen, Erik
DOI: 10.1126/science.1131007
发表时间: 2007-01-12
期刊: SCIENCE
影响因子: 56.9
作者:
Maerkl, Sebastian J.;Quake, Stephen R.
通讯作者: Quake, Stephen R.
DOI: 10.1371/journal.pbio.0060027
发表时间: 2008-02
期刊: PLoS biology
影响因子: 9.8
作者:
Li XY;MacArthur S;Bourgon R;Nix D;Pollard DA;Iyer VN;Hechmer A;Simirenko L;Stapleton M;Luengo Hendriks CL;Chu HC;Ogawa N;Inwood W;Sementchenko V;Beaton A;Weiszmann R;Celniker SE;Knowles DW;Gingeras T;Speed TP;Eisen MB;Biggin MD
通讯作者: Biggin MD
DOI: 10.1093/nar/30.5.1255
发表时间: 2002-03-01
影响因子: 14.9
作者:
Bulyk, ML;Johnson, PLF;Church, GM
通讯作者: Church, GM
DOI: 10.1371/journal.pcbi.1000156
发表时间: 2008-08-01
影响因子: 4.3
作者:
Siddharthan, Rahul
通讯作者: Siddharthan, Rahul