A Poisson mixture model to identify changes in RNA polymerase II binding quantity using high-throughput sequencing technology.

A Poisson mixture model to identify changes in RNA polymerase II binding quantity using high-throughput sequencing technology.
复制标题

DOI:
10.1186/1471-2164-9-s2-s23
复制
发表时间:
2008-09-16
期刊:
影响因子:
4.4
通讯作者:
Li L
Li L
中科院分区:
生物学2区
文献类型:
--
作者:
Feng W;Liu Y;Wu J;Nephew KP;Huang TH;Li L

文献摘要

被引文献

相似文献

我们提出了一种基于混合模型的分析,以确定RNA聚合酶II(POL II)在转录区域的分布差异,使用CHIP-SEQ(大规模平行测序技术后的染色质免疫沉淀)进行测量。统计模型假设每个基因组区域内包含的POLII靶向序列的数量服从泊松分布。然后,利用经验方法和用于估计和推断的期望最大化(EM)算法,建立了一个泊松混合模型来区分转录区的Pol II结合变化。为了在M步中获得全局最大值,实现了一种粒子群优化算法。我们将该模型应用于激素依赖的MCF7乳腺癌细胞和抗雌激素耐药的MCF7乳腺癌细胞在17β-雌二醇(E2)治疗前后产生的POL II结合数据。我们发现,在激素依赖的细胞中,约9.9%(2527)的基因在E2处理后与Pol II的结合发生了显著变化。然而,在E2处理的抗雌激素耐药细胞中,只有0.7%(172)基因显示出显著的POL II结合变化。这些结果表明,泊松混合模型可以用于芯片序列数据的分析。
We present a mixture model-based analysis for identifying differences in the distribution of RNA polymerase II (Pol II) in transcribed regions, measured using ChIP-seq (chromatin immunoprecipitation following massively parallel sequencing technology). The statistical model assumes that the number of Pol II-targeted sequences contained within each genomic region follows a Poisson distribution. A Poisson mixture model was then developed to distinguish Pol II binding changes in transcribed region using an empirical approach and an expectation-maximization (EM) algorithm developed for estimation and inference. In order to achieve a global maximum in the M-step, a particle swarm optimization (PSO) was implemented. We applied this model to Pol II binding data generated from hormone-dependent MCF7 breast cancer cells and antiestrogen-resistant MCF7 breast cancer cells before and after treatment with 17β-estradiol (E2). We determined that in the hormone-dependent cells, ~9.9% (2527) genes showed significant changes in Pol II binding after E2 treatment. However, only ~0.7% (172) genes displayed significant Pol II binding changes in E2-treated antiestrogen-resistant cells. These results show that a Poisson mixture model can be used to analyze ChIP-seq data.