On the use of resampling tests for evaluating statistical significance of binding-site co-occurrence

On the use of resampling tests for evaluating statistical significance of binding-site co-occurrence
复制标题

DOI:
10.1186/1471-2105-11-359
复制
发表时间:
2010-06-30
期刊:
影响因子:
3
通讯作者:
Russell, Steven
Russell, Steven
中科院分区:
生物学4区
文献类型:
--
作者:
Huen, David S.;Russell, Steven

文献摘要

被引文献

相似文献

背景:在真核生物中,大多数DNA结合蛋白作为大的效应复合体发挥作用。这些复合体的存在在高通量全基因组分析中通过不同复杂成分的结合位点的共现而被揭示。重抽样检验是评估表观共现的统计学意义的一种方法。结果:我们研究了两种重抽样方法来评价结合位点共现的统计意义。置换检验法被发现产生了过于有利的p值,而独立重抽样方法则具有相反的效果,在实际中几乎没有用。我们开发了一种新的、务实设计的混合方法,当应用于Polycomb/Trithorax研究的实验结果时,产生的p值与该研究的结果一致。我们将我们的研究扩展到Haiminen等人开发的FL方法,该方法从数据集中的所有结合位点推导出其零分布,并表明该方法为一对因子计算的p值可以取决于该数据集中包括哪些其他因子。我们的混合方法和FL方法似乎都对共生现象的统计意义产生了可信的估计,尽管我们的混合方法在应用于Polycomb/Trithorax数据集时更为保守。结论:我们提出了一种新的基于重采样的共现重要性测试方法,并在大型实验数据集上证明了它的性能与现有方法相同或更好。我们相信,它可以有效地应用于来自高通量全基因组技术的数据,如芯片或DAMID。本文附带了实现我们的方法的Coocate包。
Background: In eukaryotes, most DNA-binding proteins exert their action as members of large effector complexes. The presence of these complexes are revealed in high-throughput genome-wide assays by the co-occurrence of the binding sites of different complex components. Resampling tests are one route by which the statistical significance of apparent co-occurrence can be assessed.Results: We have investigated two resampling approaches for evaluating the statistical significance of binding-site co-occurrence. The permutation test approach was found to yield overly favourable p-values while the independent resampling approach had the opposite effect and is of little use in practical terms. We have developed a new, pragmatically-devised hybrid approach that, when applied to the experimental results of an Polycomb/Trithorax study, yielded p-values consistent with the findings of that study. We extended our investigations to the FL method developed by Haiminen et al, which derives its null distribution from all binding sites within a dataset, and show that the p-value computed for a pair of factors by this method can depend on which other factors are included in that dataset. Both our hybrid method and the FL method appeared to yield plausible estimates of the statistical significance of co-occurrences although our hybrid method was more conservative when applied to the Polycomb/Trithorax dataset. A high-performance parallelized implementation of the hybrid method is available.Conclusions: We propose a new resampling-based co-occurrence significance test and demonstrate that it performs as well as or better than existing methods on a large experimentally-derived dataset. We believe it can be usefully applied to data from high-throughput genome-wide techniques such as ChIP-chip or DamID. The Cooccur package, which implements our approach, accompanies this paper.