Finding sequence motifs with Bayesian models incorporating positional information: an application to transcription factor binding sites.

Finding sequence motifs with Bayesian models incorporating positional information: an application to transcription factor binding sites.
复制标题

DOI:
10.1186/1471-2105-9-262
复制
发表时间:
2008-06-04
期刊:
影响因子:
3
通讯作者:
Spouge JL
Spouge JL
中科院分区:
生物学4区
文献类型:
--
作者:
Kim NK;Tharakaraman K;Mariño-Ramírez L;Spouge JL

文献摘要

参考文献

被引文献

相似文献

生物活性序列基序通常具有相对于基因组标记的位置偏好。例如,许多已知的转录因子结合位点(TFBSs)发生在转录起始位点(TSS)上游的区间[- 300,0]个碱基内。虽然一些识别序列基序的程序利用了位置信息,但大多数程序只是隐式地和用特殊的方法对序列基序进行建模,使得它们不适合一般的基序搜索。a - glam是一个用户友好的识别序列基序的计算机程序,现在包含了一个贝叶斯模型,系统地结合了序列和位置信息。a - glam在两个人类TFBS数据集上比较了有和没有位置信息的预测结果,每个数据集都包含与已知TSS上游区间[-2000,0]个碱基相对应的序列。严格的统计分析表明,位置信息显著提高了序列基序的预测,广泛的交叉验证研究表明,A- glam模型对其参数的轻微错配具有鲁棒性。正如预期的那样,当数据集中的序列被依次截断到[- 1000,0]、[- 500,0]和[- 250,0]区间时,位置信息对基序预测的辅助作用越来越小,但对基序预测的影响并不显著。虽然在寻找具有位置偏好的生物活性基序时,序列截断是一种可行的策略,但概率模型(合理使用)通常提供更优越和更稳健的策略,特别是当序列基序的位置偏好没有很好地表征时。
Biologically active sequence motifs often have positional preferences with respect to a genomic landmark. For example, many known transcription factor binding sites (TFBSs) occur within an interval [-300, 0] bases upstream of a transcription start site (TSS). Although some programs for identifying sequence motifs exploit positional information, most of them model it only implicitly and with ad hoc methods, making them unsuitable for general motif searches. A-GLAM, a user-friendly computer program for identifying sequence motifs, now incorporates a Bayesian model systematically combining sequence and positional information. A-GLAM's predictions with and without positional information were compared on two human TFBS datasets, each containing sequences corresponding to the interval [-2000, 0] bases upstream of a known TSS. A rigorous statistical analysis showed that positional information significantly improved the prediction of sequence motifs, and an extensive cross-validation study showed that A-GLAM's model was robust against mild misspecification of its parameters. As expected, when sequences in the datasets were successively truncated to the intervals [-1000, 0], [-500, 0] and [-250, 0], positional information aided motif prediction less and less, but never hurt it significantly. Although sequence truncation is a viable strategy when searching for biologically active motifs with a positional preference, a probabilistic model (used reasonably) generally provides a superior and more robust strategy, particularly when the sequence motifs' positional preferences are not well characterized.
DOI: 10.1186/1471-2105-7-396
发表时间: 2006-08-31
期刊: BMC bioinformatics
影响因子: 3
作者:
Defrance M;Touzet H
通讯作者: Touzet H
DOI: 10.1093/nar/gkh169
发表时间: 2004-01-01
影响因子: 14.9
作者:
Frith, MC;Hansen, U;Weng, ZP
通讯作者: Weng, ZP
DOI: 10.1007/bf00993379
发表时间: 1995-10-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
BAILEY, TL;ELKAN, C
通讯作者: ELKAN, C
TransFac及其模块移植:真核生物中的转录基因调节。
DOI: 10.1093/nar/gkj143
发表时间: 2006-01-01
影响因子: 14.9
作者:
Matys V;Kel-Margoulis OV;Fricke E;Liebich I;Land S;Barre-Dirrie A;Reuter I;Chekmenev D;Krull M;Hornischer K;Voss N;Stegmaier P;Lewicki-Potapov B;Saxel H;Kel AE;Wingender E
通讯作者: Wingender E
DOI: 10.1093/nar/gkh246
发表时间: 2004-02-01
影响因子: 14.9
作者:
Mariño-Ramírez, L;Spouge, JL;Landsman, D
通讯作者: Landsman, D