De novo detection of differentially bound regions for ChIP-seq data using peaks and windows: controlling error rates correctly

De novo detection of differentially bound regions for ChIP-seq data using peaks and windows: controlling error rates correctly
复制标题

DOI:
10.1093/nar/gku351
复制
发表时间:
2014-01-01
影响因子:
14.9
通讯作者:
Smyth, Gordon K.
Smyth, Gordon K.
中科院分区:
生物学2区
文献类型:
--
作者:
Lun, Aaron T. L.;Smyth, Gordon K.

文献摘要

被引文献

相似文献

ChIP-seq实验的一个共同目的是确定不同条件下蛋白质结合模式的变化,即差异结合。当感兴趣的区域事先不知道时,已经开发了许多基于峰值和窗口的策略来检测差异结合。然而,在应用这些方法时,需要仔细考虑误差控制。基于峰值的方法使用相同的数据集来定义峰值并检测差异绑定。如果操作不当,这可能导致失去类型I错误控制。对于基于窗口的方法,控制所有检测窗口的错误发现率并不能保证控制所有检测区域。将前者误解为后者可能会导致意想不到的自由。本文提出了几种解决方案来维护这些从头计数策略的错误控制。对于基于峰值的方法,应该在统计分析之前对池库执行峰值调用。对于基于窗口的方法,提出了一种使用Simes方法的混合方法来保持对跨区域错误发现率的控制。更一般地说,使用一系列模拟和真实数据集探讨了基于峰值和基于窗口的策略的相对优势。这两种策略的实现也优于现有的差分绑定分析程序。
A common aim in ChIP-seq experiments is to identify changes in protein binding patterns between conditions, i.e. differential binding. A number of peak-and window-based strategies have been developed to detect differential binding when the regions of interest are not known in advance. However, careful consideration of error control is needed when applying these methods. Peak-based approaches use the same data set to define peaks and to detect differential binding. Done improperly, this can result in loss of type I error control. For window-based methods, controlling the false discovery rate over all detected windows does not guarantee control across all detected regions. Misinterpreting the former as the latter can result in unexpected liberalness. Here, several solutions are presented to maintain error control for these de novo counting strategies. For peak-based methods, peak calling should be performed on pooled libraries prior to the statistical analysis. For window-based methods, a hybrid approach using Simes' method is proposed to maintain control of the false discovery rate across regions. More generally, the relative advantages of peak-and window-based strategies are explored using a range of simulated and real data sets. Implementations of both strategies also compare favourably to existing programs for differential binding analyses.