Normalization, bias correction, and peak calling for ChIP-seq

Normalization, bias correction, and peak calling for ChIP-seq
复制标题

DOI:
10.1515/1544-6115.1750
复制
发表时间:
2012-01-01
影响因子:
0.9
通讯作者:
Song, Jun S.
Song, Jun S.
中科院分区:
数学4区
文献类型:
--
作者:
Diaz, Aaron;Park, Kiyoub;Song, Jun S.

文献摘要

被引文献

相似文献

下一代测序正在迅速改变我们描述细胞转录、遗传和表观遗传状态的能力。特别是,从蛋白质-DNA复合物的免疫沉淀(ChIP-seq)和甲基化DNA (MeDIP-seq)中测序DNA可以揭示蛋白质结合位点和表观遗传修饰的位置。这些方法包含许多偏差,可能会对结果数据的解释产生重大影响。检测和消除这种偏差的严格计算方法仍然缺乏。此外,多样本归一化仍然是一个重要的开放性问题。这篇理论论文通过比较62个独立的公开数据集,使用严格的统计模型和信号处理技术,系统地描述了ChIP-seq数据的偏差和特性。提出了从背景噪声中分离ChIP-seq信号的统计方法,以及对序列依赖和超声偏差的富集测试统计量进行校正。我们的方法在归一化之前有效地将读取分离为信号和背景分量,提高了信噪比。此外,大多数峰值调用者目前使用通用null模型,该模型在检测细微但真实的ChIP富集所需的灵敏度水平上特异性较低。所提出的确定细胞类型特异性零模型的方法,可以解释细胞类型特异性偏差,在给定的显著性阈值下,比当前方法能够实现更低的错误发现率。
Next-generation sequencing is rapidly transforming our ability to profile the transcriptional, genetic, and epigenetic states of a cell. In particular, sequencing DNA from the immunoprecipitation of protein-DNA complexes (ChIP-seq) and methylated DNA (MeDIP-seq) can reveal the locations of protein binding sites and epigenetic modifications. These approaches contain numerous biases which may significantly influence the interpretation of the resulting data. Rigorous computational methods for detecting and removing such biases are still lacking. Also, multi-sample normalization still remains an important open problem. This theoretical paper systematically characterizes the biases and properties of ChIP-seq data by comparing 62 separate publicly available datasets, using rigorous statistical models and signal processing techniques. Statistical methods for separating ChIP-seq signal from background noise, as well as correcting enrichment test statistics for sequence-dependent and sonication biases, are presented. Our method effectively separates reads into signal and background components prior to normalization, improving the signal-to-noise ratio. Moreover, most peak callers currently use a generic null model which suffers from low specificity at the sensitivity level requisite for detecting subtle, but true, ChIP enrichment. The proposed method of determining a cell type-specific null model, which accounts for cell type-specific biases, is shown to be capable of achieving a lower false discovery rate at a given significance threshold than current methods.