Genome-wide identification of in vivo protein-DNA binding sites from ChIP-Seq data.

Genome-wide identification of in vivo protein-DNA binding sites from ChIP-Seq data.
复制标题

DOI:
10.1093/nar/gkn488
复制
发表时间:
2008-09
影响因子:
14.9
通讯作者:
Zhao, Keji
Zhao, Keji
中科院分区:
生物学2区
文献类型:
--
作者:
Jothi, Raja;Cuddapah, Suresh;Barski, Artem;Cui, Kairong;Zhao, Keji

文献摘要

参考文献

被引文献

相似文献

染色质免疫沉淀测序(ChIP - Seq)将染色质免疫沉淀(ChIP)与超高通量大规模平行测序相结合,越来越多地用于在基因组规模上绘制体内蛋白质 - DNA相互作用图谱。通常,来自ChIP - Seq的短序列读取会被映射到参考基因组以进行进一步分析。尽管富含映射读取的基因组区域可被推断为大致的结合区域,但短读取长度(约25 - 50个核苷酸)对确定这些区域内的确切结合位点构成了挑战。在此,我们提出了SISSRs(从短序列读取中识别位点),这是一种从ChIP - Seq实验产生的短读取中精确识别结合位点的新算法。通过将SISSRs应用于三种被广泛研究且特征明确的人类转录因子(CTCF(CCCTC结合因子)、NRSF(神经元限制性沉默因子)和STAT1(信号转导子和转录激活蛋白1))的ChIP - Seq数据,证明了其敏感性和特异性。我们分别确定了26814个、5813个和73956个CTCF、NRSF和STAT1蛋白的结合位点,这比之前针对相应蛋白所推断的数量分别多出32%、299%和78%。基序分析表明,绝大多数已识别的结合位点包含先前为相应蛋白确定的共有结合序列,从而证明了SISSRs的准确性。SISSRs的敏感性和精确性有助于对ChIP - Seq数据进行进一步分析,揭示出有趣的见解,我们相信这将为设计用于绘制体内蛋白质 - DNA相互作用图谱的ChIP - Seq实验提供指导。我们还表明,结合位点处的标签密度是蛋白质 - DNA结合亲和力的良好指标,可用于区分和表征强结合位点和弱结合位点。利用标签密度作为DNA结合亲和力的指标,我们已经确定了NRSF和CTCF结合位点内对更强的DNA结合至关重要的核心残基。
ChIP-Seq, which combines chromatin immunoprecipitation (ChIP) with ultra high-throughput massively parallel sequencing, is increasingly being used for mapping protein–DNA interactions in-vivo on a genome scale. Typically, short sequence reads from ChIP-Seq are mapped to a reference genome for further analysis. Although genomic regions enriched with mapped reads could be inferred as approximate binding regions, short read lengths (∼25–50 nt) pose challenges for determining the exact binding sites within these regions. Here, we present SISSRs (Site Identification from Short Sequence Reads), a novel algorithm for precise identification of binding sites from short reads generated from ChIP-Seq experiments. The sensitivity and specificity of SISSRs are demonstrated by applying it on ChIP-Seq data for three widely studied and well-characterized human transcription factors: CTCF (CCCTC-binding factor), NRSF (neuron-restrictive silencer factor) and STAT1 (signal transducer and activator of transcription protein 1). We identified 26 814, 5813 and 73 956 binding sites for CTCF, NRSF and STAT1 proteins, respectively, which is 32, 299 and 78% more than that inferred previously for the respective proteins. Motif analysis revealed that an overwhelming majority of the identified binding sites contained the previously established consensus binding sequence for the respective proteins, thus attesting for SISSRs’ accuracy. SISSRs’ sensitivity and precision facilitated further analyses of ChIP-Seq data revealing interesting insights, which we believe will serve as guidance for designing ChIP-Seq experiments to map in vivo protein–DNA interactions. We also show that tag densities at the binding sites are a good indicator of protein–DNA binding affinity, which could be used to distinguish and characterize strong and weak binding sites. Using tag density as an indicator of DNA-binding affinity, we have identified core residues within the NRSF and CTCF binding sites that are critical for a stronger DNA binding.
DOI: 10.1523/jneurosci.0091-07.2007
发表时间: 2007-06-20
影响因子: 5.3
作者:
Otto, Stefanie J.;McCorkle, Sean R.;Mandel, Gail
通讯作者: Mandel, Gail
DOI: 10.1016/j.cell.2005.03.013
发表时间: 2005-05-20
期刊: CELL
影响因子: 64.5
作者:
Ballas, N;Grunseich, C;Mandel, G
通讯作者: Mandel, G
DOI: 10.1016/j.cell.2006.12.048
发表时间: 2007-03-23
期刊: CELL
影响因子: 64.5
作者:
Kim, Tae Hoon;Abdullaev, Ziedulla K.;Ren, Bing
通讯作者: Ren, Bing
DOI: 10.1074/jbc.m706213200
发表时间: 2007-11-16
影响因子: 4.8
作者:
Renda, Mario;Baglivo, Ilaria;Pedone, Paolo V.
通讯作者: Pedone, Paolo V.
DOI: 10.1126/science.7871435
发表时间: 1995-03-03
期刊: SCIENCE
影响因子: 56.9
作者:
SCHOENHERR, CJ;ANDERSON, DJ
通讯作者: ANDERSON, DJ