A Bayesian search for transcriptional motifs.

A Bayesian search for transcriptional motifs.
复制标题

DOI:
10.1371/journal.pone.0013897
复制
发表时间:
2010-11-18
期刊:
影响因子:
3.7
通讯作者:
Crampin EJ
Crampin EJ
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Miller AK;Print CG;Nielsen PM;Crampin EJ

文献摘要

参考文献

被引文献

相似文献

识别转录因子结合位点是理解转录调控的重要一步。一种常见的方法是使用针对特定TF的间隙对齐的、实验支持的TFBS,并且在算法上搜索相同TFBS的更多出现。TF结合特异性的最大的公开可用的数据库包含表示为位置权重矩阵(PWM)的模型。还有其他使用更复杂表示的方法,但这些方法的数据库更有限,或者无法公开获得。因此,本文重点研究了每个TF使用一个PWM进行搜索的方法。用于识别对应于特定PWM的TFBS的算法MATCHTM是可用的,但不是基于TF结合的严格统计模型,使得难以解释或调整算法的参数和输出。另一种算法MAST使用在每个位置的每个偏移处找到每个碱基的真实概率来计算TFBS存在的p值。我们开发了一个统计模型,BaSeTraM,结合TF TFBS,考虑到随机变化的基础上存在于每个位置内的TFBS。将矩阵中的计数和位点序列作为随机变量,将TFBS组合模型与背景模型相联合收割机,得到贝叶斯分类器。我们在一个包(SBaSeTraM)中实现了分类器。我们通过搜索实验酿酒酵母TF结合数据集中使用的所有探针,并将我们的预测与数据进行比较,来测试SBaSeTraM对MATCHTM的实现。我们没有发现统计学上的显着差异的算法之间的灵敏度(在固定的选择性),表明SBaSeTraM的性能至少是目前领先的算法相媲美。我们的软件可在http://wiki.github.com/A1kmm/sbasetram/building-the-tools免费获得。
Identifying transcription factor (TF) binding sites (TFBSs) is an important step towards understanding transcriptional regulation. A common approach is to use gaplessly aligned, experimentally supported TFBSs for a particular TF, and algorithmically search for more occurrences of the same TFBSs. The largest publicly available databases of TF binding specificities contain models which are represented as position weight matrices (PWM). There are other methods using more sophisticated representations, but these have more limited databases, or aren't publicly available. Therefore, this paper focuses on methods that search using one PWM per TF. An algorithm, MATCHTM, for identifying TFBSs corresponding to a particular PWM is available, but is not based on a rigorous statistical model of TF binding, making it difficult to interpret or adjust the parameters and output of the algorithm. Furthermore, there is no public description of the algorithm sufficient to exactly reproduce it. Another algorithm, MAST, computes a p-value for the presence of a TFBS using true probabilities of finding each base at each offset from that position. We developed a statistical model, BaSeTraM, for the binding of TFs to TFBSs, taking into account random variation in the base present at each position within a TFBS. Treating the counts in the matrices and the sequences of sites as random variables, we combine this TFBS composition model with a background model to obtain a Bayesian classifier. We implemented our classifier in a package (SBaSeTraM). We tested SBaSeTraM against a MATCHTM implementation by searching all probes used in an experimental Saccharomyces cerevisiae TF binding dataset, and comparing our predictions to the data. We found no statistically significant differences in sensitivity between the algorithms (at fixed selectivity), indicating that SBaSeTraM's performance is at least comparable to the leading currently available algorithm. Our software is freely available at: http://wiki.github.com/A1kmm/sbasetram/building-the-tools.
DOI: 10.1038/nbt1053
发表时间: 2005-01-01
影响因子: 46.9
作者:
Tompa, M;Li, N;Zhu, Z
通讯作者: Zhu, Z
DOI: 10.1093/nar/gkh012
发表时间: 2004-01-01
影响因子: 14.9
作者:
Sandelin, A;Alkema, W;Lenhard, B
通讯作者: Lenhard, B
DOI: 10.1371/journal.pcbi.0030061
发表时间: 2007-03-30
影响因子: 4.3
作者:
Mahony, Shaun;Auron, Philip E.;Benos, Panayiotis V.
通讯作者: Benos, Panayiotis V.
DOI: 10.1371/journal.pone.0001820
发表时间: 2008-03-26
期刊: PloS one
影响因子: 3.7
作者:
Lähdesmäki H;Rust AG;Shmulevich I
通讯作者: Shmulevich I
DOI: 10.1093/nar/gkg108
发表时间: 2003-01-01
影响因子: 14.9
作者:
Matys, V;Fricke, E;Wingender, E
通讯作者: Wingender, E