PhyloScan: identification of transcription factor binding sites using cross-species evidence

PhyloScan: identification of transcription factor binding sites using cross-species evidence
复制标题

DOI:
10.1186/1748-7188-2-1
复制
发表时间:
2007-01-23
影响因子:
1
通讯作者:
Lawrence, Charles E.
Lawrence, Charles E.
中科院分区:
生物学4区
文献类型:
--
作者:
Carmack, C. Steven;McCue, Lee Ann;Lawrence, Charles E.

文献摘要

被引文献

相似文献

背景资料:当已知特定转录因子的转录因子结合位点时,可以构建可用于扫描序列以寻找其他位点的基序模型。然而,很少有统计学意义的网站被发现时,转录因子结合位点基序模型被用来扫描基因组规模database.Methods:我们已经开发了一种扫描算法,PhyloScan,它结合了证据从匹配的网站中发现的orthopathy数据从几个相关的物种与证据从多个网站内的一个基因间区域,以更好地检测调节子。正向序列数据可以是多重比对的、未比对的或比对和未比对的组合。在对齐的数据中,PhyloScan在统计上考虑了对对齐贡献数据的物种的系统发育依赖性,并且在未对齐的数据中,假设物种的系统发育独立性,结合了位点的证据。直接计算的基因预测的统计学意义,而不采用training sets.Results:在我们的方法的合成数据上模拟七个肠杆菌目,四个弧菌目,和三个巴斯德菌目物种的测试,PhyloScan产生更好的灵敏度和特异性比猴子,先进的扫描方法,也搜索基因组的转录因子结合位点使用系统发育信息。应用该算法从7个肠杆菌属物种的真实的序列数据确定新的CRP和PurR转录因子结合位点,从而为这些转录因子提供了几个新的潜在位点。这些位点使得有针对性的实验验证,从而进一步描绘了大肠杆菌中的CRP和PurR调节子。coli.Conclusion:更好的灵敏度和特异性可以通过(1)使用混合的可重复和不可重复序列数据和(2)组合来自基因间区域内多个位点的证据来实现。
Background: When transcription factor binding sites are known for a particular transcription factor, it is possible to construct a motif model that can be used to scan sequences for additional sites. However, few statistically significant sites are revealed when a transcription factor binding site motif model is used to scan a genome-scale database.Methods: We have developed a scanning algorithm, PhyloScan, which combines evidence from matching sites found in orthologous data from several related species with evidence from multiple sites within an intergenic region, to better detect regulons. The orthologous sequence data may be multiply aligned, unaligned, or a combination of aligned and unaligned. In aligned data, PhyloScan statistically accounts for the phylogenetic dependence of the species contributing data to the alignment and, in unaligned data, the evidence for sites is combined assuming phylogenetic independence of the species. The statistical significance of the gene predictions is calculated directly, without employing training sets.Results: In a test of our methodology on synthetic data modeled on seven Enterobacteriales, four Vibrionales, and three Pasteurellales species, PhyloScan produces better sensitivity and specificity than MONKEY, an advanced scanning approach that also searches a genome for transcription factor binding sites using phylogenetic information. The application of the algorithm to real sequence data from seven Enterobacteriales species identifies novel Crp and PurR transcription factor binding sites, thus providing several new potential sites for these transcription factors. These sites enable targeted experimental validation and thus further delineation of the Crp and PurR regulons in E. coli.Conclusion: Better sensitivity and specificity can be achieved through a combination of (1) using mixed alignable and non-alignable sequence data and (2) combining evidence from multiple sites within an intergenic region.