Discovering transcription factor binding sites in highly repetitive regions of genomes with multi-read analysis of ChIP-Seq data.

Discovering transcription factor binding sites in highly repetitive regions of genomes with multi-read analysis of ChIP-Seq data.
复制标题

DOI:
10.1371/journal.pcbi.1002111
复制
发表时间:
2011-07
影响因子:
4.3
通讯作者:
Keleş S
Keleş S
中科院分区:
生物学2区
文献类型:
--
作者:
Chung D;Kuan PF;Li B;Sanalkumar R;Liang K;Bresnick EH;Dewey C;Keleş S

文献摘要

参考文献

被引文献

相似文献

染色质免疫沉淀和高通量测序(CHIP-SEQ)正在迅速取代染色质免疫沉淀和全基因组平铺阵列分析(CHIP-CHIP),成为定位转录因子结合位点和染色质修饰的首选方法。分析芯片序列数据的最新技术依赖于仅使用唯一映射到相关参考基因组的读取(单一读取)。这可能会导致高达30%的可对齐读取被遗漏。我们描述了一种利用映射到参考基因组上多个位置的读出(多读)的一般方法。我们的方法基于使用加权对齐方案将多次读取分配为分数计数。使用人类STAT1和小鼠GATA1芯片序列数据集,我们证明了多读入显著增加了测序深度,导致检测到以其他方式无法用单读识别的新峰,并改进了对可映射区中峰的检测。我们通过计算实验研究了只有通过多次读取才能检测到的峰的各种全基因组特征。总体而言,来自多读分析的峰与由单读鉴定的峰具有相似的特征,除了它们中的大多数驻留在分段复制中。我们进一步通过独立的定量实时芯片分析验证了一些GATA1多读峰,并发现了GATA1的新的靶基因。这些计算和实验结果表明,多读取对于用芯片序列实验研究基因组高重复区域中的转录因子结合至关重要。对基因组重复区域进行实验注释是一项具有挑战性的任务。染色质免疫沉淀和高通量测序(CHIP-SEQ)为从转录因子结合的角度表征基因组的重复区域提供了有价值的数据。虽然芯片序列技术已经成熟,但可用的芯片序列分析方法和软件依赖于丢弃映射到参考基因组上多个位置的序列读取(多读取),从而错过了评估转录因子与基因组高度重复区域结合的机会。我们开发了一种在芯片序列分析中考虑多次读取的计算算法。我们通过计算实验表明,多读取导致测序深度的显着增加和结合区域的识别,否则当只使用唯一映射到参考基因组的读出(单读取)时,这些结合区域是无法识别的。特别是,我们发现识别的结合区数量可以增加高达36%。我们通过独立的实时芯片定量验证结合区域来支持我们的计算预测,只有当多次读取被纳入到小鼠GATA1芯片序列实验的分析中时才能识别。
Chromatin immunoprecipitation followed by high-throughput sequencing (ChIP-seq) is rapidly replacing chromatin immunoprecipitation combined with genome-wide tiling array analysis (ChIP-chip) as the preferred approach for mapping transcription-factor binding sites and chromatin modifications. The state of the art for analyzing ChIP-seq data relies on using only reads that map uniquely to a relevant reference genome (uni-reads). This can lead to the omission of up to 30% of alignable reads. We describe a general approach for utilizing reads that map to multiple locations on the reference genome (multi-reads). Our approach is based on allocating multi-reads as fractional counts using a weighted alignment scheme. Using human STAT1 and mouse GATA1 ChIP-seq datasets, we illustrate that incorporation of multi-reads significantly increases sequencing depths, leads to detection of novel peaks that are not otherwise identifiable with uni-reads, and improves detection of peaks in mappable regions. We investigate various genome-wide characteristics of peaks detected only by utilization of multi-reads via computational experiments. Overall, peaks from multi-read analysis have similar characteristics to peaks that are identified by uni-reads except that the majority of them reside in segmental duplications. We further validate a number of GATA1 multi-read only peaks by independent quantitative real-time ChIP analysis and identify novel target genes of GATA1. These computational and experimental results establish that multi-reads can be of critical importance for studying transcription factor binding in highly repetitive regions of genomes with ChIP-seq experiments. Annotating repetitive regions of genomes experimentally is a challenging task. Chromatin immunoprecipitation followed by high-throughput sequencing (ChIP-seq) provides valuable data for characterizing repetitive regions of genomes in terms of transcription factor binding. Although ChIP-seq technology has been maturing, available ChIP-seq analysis methods and software rely on discarding sequence reads that map to multiple locations on the reference genome (multi-reads), thereby generating a missed opportunity for assessing transcription factor binding to highly repetitive regions of genomes. We develop a computational algorithm that takes multi-reads into account in ChIP-seq analysis. We show with computational experiments that multi-reads lead to significant increase in sequencing depths and identification of binding regions that are otherwise not identifiable when only reads that uniquely map to the reference genome (uni-reads) are used. In particular, we show that the number of binding regions identified can increase up to 36%. We support our computational predictions with independent quantitative real-time ChIP validation of binding regions identified only when multi-reads are incorporated in the analysis of a mouse GATA1 ChIP-seq experiment.
DOI: 10.1016/s0092-8674(04)00127-8
发表时间: 2004-02-20
期刊: CELL
影响因子: 64.5
作者:
Cawley, S;Bekiranov, S;Gingeras, TR
通讯作者: Gingeras, TR
DOI: 10.1186/gb-2010-11-6-r69
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Day DS;Luquette LJ;Park PJ;Kharchenko PV
通讯作者: Kharchenko PV
DOI: 10.1093/nar/gkp1012
发表时间: 2010-01
影响因子: 14.9
作者:
Blahnik KR;Dou L;O'Geen H;McPhillips T;Xu X;Cao AR;Iyengar S;Nicolet CM;Ludäscher B;Korf I;Farnham PJ
通讯作者: Farnham PJ
DOI: 10.1038/nbt.1505
发表时间: 2008-11
影响因子: 46.9
作者:
Ji, Hongkai;Jiang, Hui;Ma, Wenxiu;Johnson, David S.;Myers, Richard M.;Wong, Wing H.
通讯作者: Wong, Wing H.
DOI: 10.1016/j.molcel.2009.11.001
发表时间: 2009-11-25
期刊: Molecular cell
影响因子: 16
作者:
Fujiwara T;O'Geen H;Keles S;Blahnik K;Linnemann AK;Kang YA;Choi K;Farnham PJ;Bresnick EH
通讯作者: Bresnick EH