Discovering transcription factor binding sites in highly repetitive regions of genomes with multi-read analysis of ChIP-Seq data.
Discovering transcription factor binding sites in highly repetitive regions of genomes with multi-read analysis of ChIP-Seq data.
复制标题
DOI:
10.1371/journal.pcbi.1002111
复制
发表时间:
2011-07
影响因子:
4.3
通讯作者:
Keleş S
中科院分区:
文献类型:
--
作者:
Chung D;Kuan PF;Li B;Sanalkumar R;Liang K;Bresnick EH;Dewey C;Keleş S
Chromatin immunoprecipitation followed by high-throughput sequencing (ChIP-seq) is rapidly replacing chromatin immunoprecipitation combined with genome-wide tiling array analysis (ChIP-chip) as the preferred approach for mapping transcription-factor binding sites and chromatin modifications. The state of the art for analyzing ChIP-seq data relies on using only reads that map uniquely to a relevant reference genome (uni-reads). This can lead to the omission of up to 30% of alignable reads. We describe a general approach for utilizing reads that map to multiple locations on the reference genome (multi-reads). Our approach is based on allocating multi-reads as fractional counts using a weighted alignment scheme. Using human STAT1 and mouse GATA1 ChIP-seq datasets, we illustrate that incorporation of multi-reads significantly increases sequencing depths, leads to detection of novel peaks that are not otherwise identifiable with uni-reads, and improves detection of peaks in mappable regions. We investigate various genome-wide characteristics of peaks detected only by utilization of multi-reads via computational experiments. Overall, peaks from multi-read analysis have similar characteristics to peaks that are identified by uni-reads except that the majority of them reside in segmental duplications. We further validate a number of GATA1 multi-read only peaks by independent quantitative real-time ChIP analysis and identify novel target genes of GATA1. These computational and experimental results establish that multi-reads can be of critical importance for studying transcription factor binding in highly repetitive regions of genomes with ChIP-seq experiments. Annotating repetitive regions of genomes experimentally is a challenging task. Chromatin immunoprecipitation followed by high-throughput sequencing (ChIP-seq) provides valuable data for characterizing repetitive regions of genomes in terms of transcription factor binding. Although ChIP-seq technology has been maturing, available ChIP-seq analysis methods and software rely on discarding sequence reads that map to multiple locations on the reference genome (multi-reads), thereby generating a missed opportunity for assessing transcription factor binding to highly repetitive regions of genomes. We develop a computational algorithm that takes multi-reads into account in ChIP-seq analysis. We show with computational experiments that multi-reads lead to significant increase in sequencing depths and identification of binding regions that are otherwise not identifiable when only reads that uniquely map to the reference genome (uni-reads) are used. In particular, we show that the number of binding regions identified can increase up to 36%. We support our computational predictions with independent quantitative real-time ChIP validation of binding regions identified only when multi-reads are incorporated in the analysis of a mouse GATA1 ChIP-seq experiment.
登录
查看更多内容
影响因子:
64.5
作者:
Cawley, S;Bekiranov, S;Gingeras, TR
通讯作者:
Gingeras, TR
影响因子:
12.3
作者:
Day DS;Luquette LJ;Park PJ;Kharchenko PV
通讯作者:
Kharchenko PV
影响因子:
14.9
作者:
Blahnik KR;Dou L;O'Geen H;McPhillips T;Xu X;Cao AR;Iyengar S;Nicolet CM;Ludäscher B;Korf I;Farnham PJ
通讯作者:
Farnham PJ
影响因子:
46.9
作者:
Ji, Hongkai;Jiang, Hui;Ma, Wenxiu;Johnson, David S.;Myers, Richard M.;Wong, Wing H.
通讯作者:
Wong, Wing H.
影响因子:
16
作者:
Fujiwara T;O'Geen H;Keles S;Blahnik K;Linnemann AK;Kang YA;Choi K;Farnham PJ;Bresnick EH
通讯作者:
Bresnick EH