AREM: Aligning Short Reads from ChIP-Sequencing by Expectation Maximization

AREM: Aligning Short Reads from ChIP-Sequencing by Expectation Maximization
复制标题

DOI:
10.1089/cmb.2011.0185
复制
发表时间:
2011-11-01
影响因子:
1.7
通讯作者:
Xie, Xiaohui
Xie, Xiaohui
中科院分区:
生物学4区
文献类型:
--
作者:
Newkirk, Daniel;Biesinger, Jacob;Xie, Xiaohui

文献摘要

被引文献

相似文献

高通量测序与染色质免疫沉淀(ChIP-Seq)偶联被广泛用于表征转录因子、辅因子、染色质修饰剂和其他DNA结合蛋白的全基因组结合模式。ChIP-Seq数据分析的关键步骤是将来自高通量测序的短读段映射到参考基因组,并识别富含短读段的峰区域。尽管已经提出了几种用于ChIP-Seq分析的方法,但大多数现有方法仅考虑可以唯一放置在参考基因组中的读取,因此检测位于重复序列内的峰的能力较低。在这里,我们介绍了一种利用所有读取的ChIP-Seq数据分析的概率方法,提供了一个真正的全基因组结合模式视图。使用对应于K富集区域和空基因组背景的混合物模型对读数进行建模。我们使用最大似然估计富集区域的位置,并实现期望最大化(E-M)算法,称为AREM(通过期望最大化对齐读取),以更新每个读取到不同基因组位置的对齐概率。我们应用该算法来识别两种蛋白质的全基因组结合事件:Rad 21,一种粘着蛋白的组分和参与染色单体粘着的关键因子,以及Srebp-1,一种对脂质/胆固醇稳态重要的转录因子。使用AREM,我们能够以高置信度鉴定小鼠基因组中的19,935个Rad 21峰和1,748个Srebp-1峰,包括仅使用唯一映射的读数遗漏的1,517个(7.6%)Rad 21峰和227个(13%)Srebp-1峰。我们的算法的开源实现可在http://sourceforge.net/projects/arem上获得。
High-throughput sequencing coupled to chromatin immunoprecipitation (ChIP-Seq) is widely used in characterizing genome-wide binding patterns of transcription factors, co-factors, chromatin modifiers, and other DNA binding proteins. A key step in ChIP-Seq data analysis is to map short reads from high-throughput sequencing to a reference genome and identify peak regions enriched with short reads. Although several methods have been proposed for ChIP-Seq analysis, most existing methods only consider reads that can be uniquely placed in the reference genome, and therefore have low power for detecting peaks located within repeat sequences. Here, we introduce a probabilistic approach for ChIP-Seq data analysis that utilizes all reads, providing a truly genome-wide view of binding patterns. Reads are modeled using a mixture model corresponding to K enriched regions and a null genomic background. We use maximum likelihood to estimate the locations of the enriched regions, and implement an expectation-maximization (E-M) algorithm, called AREM (aligning reads by expectation maximization), to update the alignment probabilities of each read to different genomic locations. We apply the algorithm to identify genome-wide binding events of two proteins: Rad21, a component of cohesin and a key factor involved in chromatid cohesion, and Srebp-1, a transcription factor important for lipid/cholesterol homeostasis. Using AREM, we were able to identify 19,935 Rad21 peaks and 1,748 Srebp-1 peaks in the mouse genome with high confidence, including 1,517 (7.6%) Rad21 peaks and 227 (13%) Srebp-1 peaks that were missed using only uniquely mapped reads. The open source implementation of our algorithm is available at http://sourceforge.net/projects/arem.