Accounting for immunoprecipitation efficiencies in the statistical analysis of ChIP-seq data.

Accounting for immunoprecipitation efficiencies in the statistical analysis of ChIP-seq data.
复制标题

DOI:
10.1186/1471-2105-14-169
复制
发表时间:
2013-05-30
期刊:
影响因子:
3
通讯作者:
't Hoen PA
't Hoen PA
中科院分区:
生物学4区
文献类型:
--
作者:
Bao Y;Vinciotti V;Wit E;'t Hoen PA

文献摘要

参考文献

被引文献

相似文献

不同抗体之间以及使用相同抗体的重复实验之间的免疫沉淀 (IP) 效率可能会有很大差异。这些差异对 ChIP-seq 数据的质量有很大影响:与效率较低的实验相比,更高效的实验必然会导致更高的信号背景比,因此会产生明显更多的富集区域。在本文中,我们展示了如何在 ChIP-seq 数据的联合统计建模中明确考虑 IP 效率。我们将潜在混合物模型拟合到来自两个实验室的两种蛋白质的八次实验中,其中两种蛋白质使用了不同的抗体。我们使用模型参数来估计各个实验的效率,并发现不同实验室以及同一实验室的技术复制之间的效率明显不同。当我们考虑 ChIP 效率时,我们发现在相同的错误发现率下,效率较高的实验中比效率较低的实验中结合的区域更多。模型中还可以包含跨实验相同数量的结合位点的先验知识,以便更可靠地检测两种不同蛋白质之间的差异结合区域。我们提出了一种统计模型,用于从多个 ChIP-seq 数据集中检测富集和差异结合区域。我们提出的框架明确考虑了 ChIP-seq 数据中的 IP 效率,并允许对不同蛋白质进行联合(而不是单独)复制和实验建模,从而得出更可靠的生物学结论。
ImmunoPrecipitation (IP) efficiencies may vary largely between different antibodies and between repeated experiments with the same antibody. These differences have a large impact on the quality of ChIP-seq data: a more efficient experiment will necessarily lead to a higher signal to background ratio, and therefore to an apparent larger number of enriched regions, compared to a less efficient experiment. In this paper, we show how IP efficiencies can be explicitly accounted for in the joint statistical modelling of ChIP-seq data. We fit a latent mixture model to eight experiments on two proteins, from two laboratories where different antibodies are used for the two proteins. We use the model parameters to estimate the efficiencies of individual experiments, and find that these are clearly different for the different laboratories, and amongst technical replicates from the same lab. When we account for ChIP efficiency, we find more regions bound in the more efficient experiments than in the less efficient ones, at the same false discovery rate. A priori knowledge of the same number of binding sites across experiments can also be included in the model for a more robust detection of differentially bound regions among two different proteins. We propose a statistical model for the detection of enriched and differentially bound regions from multiple ChIP-seq data sets. The framework that we present accounts explicitly for IP efficiencies in ChIP-seq data, and allows to model jointly, rather than individually, replicates and experiments from different proteins, leading to more robust biological conclusions.
DOI: 10.1093/bioinformatics/btq669
发表时间: 2011-02-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lan X;Bonneville R;Apostolos J;Wu W;Jin VX
通讯作者: Jin VX
DOI: 10.1186/1471-2105-9-523
发表时间: 2008-12-05
期刊: BMC bioinformatics
影响因子: 3
作者:
Nix DA;Courdy SJ;Boucher KM
通讯作者: Boucher KM
DOI: 10.1093/nar/gkp1012
发表时间: 2010-01
影响因子: 14.9
作者:
Blahnik KR;Dou L;O'Geen H;McPhillips T;Xu X;Cao AR;Iyengar S;Nicolet CM;Ludäscher B;Korf I;Farnham PJ
通讯作者: Farnham PJ
DOI: 10.1093/nar/gks048
发表时间: 2012-05
影响因子: 14.9
作者:
Micsinai M;Parisi F;Strino F;Asp P;Dynlacht BD;Kluger Y
通讯作者: Kluger Y
DOI: 10.1038/nbt.1505
发表时间: 2008-11
影响因子: 46.9
作者:
Ji, Hongkai;Jiang, Hui;Ma, Wenxiu;Johnson, David S.;Myers, Richard M.;Wong, Wing H.
通讯作者: Wong, Wing H.