dPeak: high resolution identification of transcription factor binding sites from PET and SET ChIP-Seq data.

dPeak: high resolution identification of transcription factor binding sites from PET and SET ChIP-Seq data.
复制标题

DOI:
10.1371/journal.pcbi.1003246
复制
发表时间:
2013
影响因子:
4.3
通讯作者:
Keleş S
Keleş S
中科院分区:
生物学2区
文献类型:
--
作者:
Chung D;Park D;Myers K;Grass J;Kiley P;Landick R;Keleş S

文献摘要

参考文献

被引文献

相似文献

染色质免疫沉淀和高通量测序(ChIP-Seq)已经成功地用于许多模式生物和人类的转录因子结合位点、组蛋白修饰和核小体占用的全基因组分析。由于原核生物紧凑的基因组包含许多仅由几个碱基对分开的结合位点,ChIP-Seq在该领域的应用尚未充分发挥其潜力。ChIP-Seq在原核生物基因组中的应用受到了进一步的阻碍,因为对ChIP-Seq进行了充分研究的数据分析方法并没有产生解码附近结合事件位置所需的分辨率。我们生成了大肠杆菌因子的单端标签(SET)和成对标签(PET) ChIP-Seq数据。这些数据集的直接比较表明,尽管PET分析能够更高分辨率地识别结合事件,但标准ChIP-Seq分析方法无法利用数据的PET特异性特征。为了解决这个问题,我们开发了dPeak作为高分辨率结合位点识别(反卷积)算法。dPeak实现了一个概率模型,准确地描述了SET和PET分析的ChIP-Seq数据生成过程。对于SET数据,dPeak优于或与最先进的高分辨率ChIP-Seq峰值反卷积算法(如PICS, GPS和GEM)相当。当与PET数据相结合时,dPeak明显优于当前任何最先进的基于set的分析方法。来自PET ChIP-Seq数据的dPeak预测子集的实验验证表明,dPeak可以以高达10的分辨率估计结合事件的位置。dPeak对大肠杆菌在好氧和厌氧条件下ChIP-Seq数据的应用揭示了位于不同位置的启动子,进一步说明了ChIP-Seq数据高分辨率分析的重要性。染色质免疫沉淀后高通量测序(ChIP-Seq)被广泛用于研究体内蛋白质- dna全基因组相互作用。目前最先进的ChIP-Seq协议使用单端标签(SET)测定,仅对文库中DNA片段的末端进行测序。尽管PET测序通常用于下一代测序的其他应用,但它还没有适应ChIP-Seq。我们通过实验和计算证明,PET测序显著提高了ChIP-Seq实验的分辨率,并使ChIP-Seq应用于大肠杆菌(E. coli)等紧凑基因组。为了使用PET ChIP-Seq数据进行高效识别,我们开发了dPeak作为高分辨率结合位点识别算法。dPeak实现了SET和PET数据的概率模型,并促进了这两种数据类型的有效分析。将dPeak应用于大肠杆菌PET和SET ChIP-Seq深度测序数据,与SET测序相比,PET分辨率明显提高。
Chromatin immunoprecipitation followed by high throughput sequencing (ChIP-Seq) has been successfully used for genome-wide profiling of transcription factor binding sites, histone modifications, and nucleosome occupancy in many model organisms and humans. Because the compact genomes of prokaryotes harbor many binding sites separated by only few base pairs, applications of ChIP-Seq in this domain have not reached their full potential. Applications in prokaryotic genomes are further hampered by the fact that well studied data analysis methods for ChIP-Seq do not result in a resolution required for deciphering the locations of nearby binding events. We generated single-end tag (SET) and paired-end tag (PET) ChIP-Seq data for factor in Escherichia coli (E. coli). Direct comparison of these datasets revealed that although PET assay enables higher resolution identification of binding events, standard ChIP-Seq analysis methods are not equipped to utilize PET-specific features of the data. To address this problem, we developed dPeak as a high resolution binding site identification (deconvolution) algorithm. dPeak implements a probabilistic model that accurately describes ChIP-Seq data generation process for both the SET and PET assays. For SET data, dPeak outperforms or performs comparably to the state-of-the-art high-resolution ChIP-Seq peak deconvolution algorithms such as PICS, GPS, and GEM. When coupled with PET data, dPeak significantly outperforms SET-based analysis with any of the current state-of-the-art methods. Experimental validations of a subset of dPeak predictions from PET ChIP-Seq data indicate that dPeak can estimate locations of binding events with as high as to resolution. Applications of dPeak to ChIP-Seq data in E. coli under aerobic and anaerobic conditions reveal closely located promoters that are differentially occupied and further illustrate the importance of high resolution analysis of ChIP-Seq data. Chromatin immunoprecipitation followed by high throughput sequencing (ChIP-Seq) is widely used for studying in vivo protein-DNA interactions genome-wide. Current state-of-the-art ChIP-Seq protocols utilize single-end tag (SET) assay which only sequences ends of DNA fragments in the library. Although paired-end tag (PET) sequencing is routinely used in other applications of next generation sequencing, it has not been much adapted to ChIP-Seq. We illustrate both experimentally and computationally that PET sequencing significantly improves the resolution of ChIP-Seq experiments and enables ChIP-Seq applications in compact genomes like Escherichia coli (E. coli). To enable efficient identification using PET ChIP-Seq data, we develop dPeak as a high resolution binding site identification algorithm. dPeak implements probabilistic models for both SET and PET data and facilitates efficient analysis of both data types. Applications of dPeak to deeply sequenced E. coli PET and SET ChIP-Seq data establish significantly better resolution of PET compared to SET sequencing.
DOI: 10.1038/nbt.1505
发表时间: 2008-11
影响因子: 46.9
作者:
Ji, Hongkai;Jiang, Hui;Ma, Wenxiu;Johnson, David S.;Myers, Richard M.;Wong, Wing H.
通讯作者: Wong, Wing H.
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1371/journal.pcbi.1002638
发表时间: 2012
影响因子: 4.3
作者:
Guo Y;Mahony S;Gifford DK
通讯作者: Gifford DK
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1128/jb.119.3.736-747.1974
发表时间: 1974-01-01
影响因子: 3.2
作者:
NEIDHARDT, FC;BLOCH, PL;SMITH, DF
通讯作者: SMITH, DF