Determining significance of pairwise co-occurrences of events in bursty sequences

Determining significance of pairwise co-occurrences of events in bursty sequences
复制标题

DOI:
10.1186/1471-2105-9-336
复制
发表时间:
2008-08-08
期刊:
影响因子:
3
通讯作者:
Terzi, Evimaria
Terzi, Evimaria
中科院分区:
生物学4区
文献类型:
--
作者:
Haiminen, Niina;Mannila, Heikki;Terzi, Evimaria

文献摘要

被引文献

相似文献

背景:不同类型的事件经常发生在一起的事件序列出现,即。例如,在一个实施例中,当研究DNA序列中某些转录因子(TF,类型)的潜在转录因子结合位点(TFBS,事件)时。这些事件倾向于突然发生:在一些基因组区域中有更多的基因,因此可能有更多的结合位点,而在一些可能非常长的区域中,几乎没有任何事件发生。此外,某些类型的事件可能会发生在序列中更频繁比others.Trends的共同出现的结合位点的两个或多个TF是有趣的,因为它们可能意味着在调节过程中的TF之间的合作作用。确定数值以概括两个TF之间的共现趋势可以以多种方式完成。然而,测试这些值的意义应该做一个相关的空模型,考虑到全球sequence structure.Results:我们扩展了现有的技术,已被认为是确定一对事件类型之间的同现模式的意义下不同的空模型。这些模型从非常简单的模型到考虑序列突发性的更复杂的模型。我们评估的模型和技术的合成事件序列,并在真实的数据组成的潜在的转录因子binding sites.Conclusion:我们表明,简单的空模型是不适合突发性的数据,他们产生许多假阳性。更复杂的模型在我们的实验中给出了更好的结果。我们还演示了窗口大小的影响,即,显著性结果上的最大共现距离。
Background: Event sequences where different types of events often occur close together arise, e. g., when studying potential transcription factor binding sites (TFBS, events) of certain transcription factors (TF, types) in a DNA sequence. These events tend to occur in bursts: in some genomic regions there are more genes and therefore potentially more binding sites, while in some, possibly very long regions, hardly any events occur. Also some types of events may occur in the sequence more often than others.Tendencies of co-occurrence of binding sites of two or more TFs are interesting, as they may imply a co-operative role between the TFs in regulatory processes. Determining a numerical value to summarize the tendency for co-occurrence between two TFs can be done in a number of ways. However, testing for the significance of such values should be done with respect to a relevant null model that takes into account the global sequence structure.Results: We extend the existing techniques that have been considered for determining the significance of co-occurrence patterns between a pair of event types under different null models. These models range from very simple ones to more complex models that take the burstiness of sequences into account. We evaluate the models and techniques on synthetic event sequences, and on real data consisting of potential transcription factor binding sites.Conclusion: We show that simple null models are poorly suited for bursty data, and they yield many false positives. More sophisticated models give better results in our experiments. We also demonstrate the effect of the window size, i.e., maximum co-occurrence distance, on the significance results.