Modeling ChIP sequencing in silico with applications.

Modeling ChIP sequencing in silico with applications.
复制标题

DOI:
10.1371/journal.pcbi.1000158
复制
发表时间:
2008-08-22
影响因子:
4.3
通讯作者:
Gerstein M
Gerstein M
中科院分区:
生物学2区
文献类型:
--
作者:
Zhang ZD;Rozowsky J;Snyder M;Chang J;Gerstein M

文献摘要

参考文献

被引文献

相似文献

ChIP测序(ChIP-seq)是一种新的DNA蛋白结合位点全基因组定位方法。它在功能基因组学领域引起了极大的兴奋。为了对数据进行评分并确定足够的测序深度,必须对基因组背景和结合位点进行适当的建模。为了建立解决这些问题的计算基础,我们首先进行了一项研究,以表征这种新型高通量数据的观察到的统计性质。通过将序列标签连接到簇中,我们表明在最近的一些实验中观察到的标签计数分布有两个组成部分:初始幂律分布和随后的长右尾。然后,我们在芯片上开发了ChIP-seq,一种计算方法,通过根据实际结合位点和背景基因组序列的特定假设分布将标签放置在基因组上来模拟实验结果。与目前的假设相反,我们的结果表明,背景位点和结合位点都需要具有明显的非均匀分布,以便正确地模拟观察到的ChIP-seq数据,例如,背景标签计数由gamma分布建模。在这些结果的基础上,我们通过使用更现实的基因组背景模型扩展了现有的评分方法。这使我们能够在ChIP-seq数据中以统计严谨的方式识别转录因子结合位点。ChIP-seq是染色体免疫沉淀和下一代测序的恰当结合,可在全基因组范围内鉴定体内转录因子结合位点。自出现以来,这种新方法在功能基因组学领域引起了极大的兴奋。对ChIP-seq过程进行适当的计算建模,为数据评分和确定足够的测序深度提供了计算基础,为分析ChIP-seq数据提供了计算基础。在我们的研究中,我们展示了ChIP-seq数据的特点,并提出了芯片测序,一种模拟实验结果的计算方法。根据我们的数据表征,我们观察到转录因子结合位点序列标签过度富集。我们的模拟结果表明,基因组背景和结合位点都不均匀。基于我们的模拟结果,我们提出了一个统计程序,使用更现实的基因组背景模型来识别ChIP-seq数据中的结合位点。
ChIP sequencing (ChIP-seq) is a new method for genomewide mapping of protein binding sites on DNA. It has generated much excitement in functional genomics. To score data and determine adequate sequencing depth, both the genomic background and the binding sites must be properly modeled. To develop a computational foundation to tackle these issues, we first performed a study to characterize the observed statistical nature of this new type of high-throughput data. By linking sequence tags into clusters, we show that there are two components to the distribution of tag counts observed in a number of recent experiments: an initial power-law distribution and a subsequent long right tail. Then we develop in silico ChIP-seq, a computational method to simulate the experimental outcome by placing tags onto the genome according to particular assumed distributions for the actual binding sites and for the background genomic sequence. In contrast to current assumptions, our results show that both the background and the binding sites need to have a markedly nonuniform distribution in order to correctly model the observed ChIP-seq data, with, for instance, the background tag counts modeled by a gamma distribution. On the basis of these results, we extend an existing scoring approach by using a more realistic genomic-background model. This enables us to identify transcription-factor binding sites in ChIP-seq data in a statistically rigorous fashion. ChIP-seq is an apt combination of chromosome immunoprecipitation and next-generation sequencing to identify transcription factor binding sites in vivo on the whole-genome scale. Since its advent, this new method has generated much excitement in the field of functional genomics. Proper computational modeling of the ChIP-seq process is needed for both data scoring and determination of adequate sequencing depth, as it provides the computational foundation for analyzing ChIP-seq data. In our study, we show the characteristics of ChIP-seq data and present in silico ChIP sequencing, a computational method to simulate the experimental outcome. On the basis of our data characterization, we observed transcription factor binding sites with excessive enrichment of sequence tags. Our simulation results reveal that both the genomic background and the binding sites are not uniform. On the basis of our simulation results, we propose a statistical procedure using the more realistic genomic background model to identify binding sites in ChIP-seq data.
DOI: 10.1016/j.cell.2005.10.042
发表时间: 2006-01-13
期刊: CELL
影响因子: 64.5
作者:
Hallikas, O;Palin, K;Taipale, J
通讯作者: Taipale, J
DOI: 10.1186/gb-2007-8-5-r81
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Zhang ZD;Rozowsky J;Lam HY;Du J;Snyder M;Gerstein M
通讯作者: Gerstein M
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1074/jbc.m209566200
发表时间: 2002-12-20
影响因子: 4.8
作者:
Lemasson, I;Polakowski, NJ;Nyborg, JK
通讯作者: Nyborg, JK
DOI: 10.1038/35054095
发表时间: 2001-01-25
期刊: NATURE
影响因子: 64.8
作者:
Iyer, VR;Horak, CE;Brown, PO
通讯作者: Brown, PO