Analyzing histone ChIP-seq data with a bin-based probability of being signal.

Analyzing histone ChIP-seq data with a bin-based probability of being signal.
复制标题

DOI:
10.1371/journal.pcbi.1011568
复制
发表时间:
2023-10
影响因子:
4.3
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

组蛋白ChIP-seq是绘制细胞表观基因组景观的主要方法之一,其组分在基因表达中起着关键的调节作用。分析跨数据集和细胞类型的调控元件的活性可能具有挑战性,这是由于例如不同的读取深度、ChIP效率和靶大小导致的峰位置偏移和归一化伪影。此外,在抑制性组蛋白标记中观察到的广泛富集区域通常逃避常用峰识别器的检测。在这里,我们提出了一种简单而通用的方法,用于识别ChIP-seq数据中的富集区域,该方法依赖于估计与非重叠的5 kB基因组箱拟合的伽马分布,以建立全局背景。我们使用该分布来为每个5 kB bin分配介于0和1之间的信号概率(PBS)。这种方法虽然分辨率低于典型的峰调用方法,但提供了一种直接的方法来识别富集区域并比较多个数据集之间的富集,方法是将数据转换为通用标准化的值,并且可以很容易地可视化并与下游分析方法集成。我们展示了PBS的广泛和狭窄的组蛋白标记的应用,并提供了几个生物学的见解,可以通过整合PBS分数与下游数据类型收集的插图。组蛋白修饰是基因表达的表观遗传调控的关键。它们的基因组分布最常使用ChIP-seq测量,ChIP-seq是一种下一代测序方法,其导致在具有修饰的组蛋白的区域附近的测序读数堆积。由于缺乏用于解释信号强度的明确标准、组蛋白位置的固有可变性以及检测与背景相比仅适度富集的非常广泛的区域的困难,比较不同细胞背景中组蛋白修饰谱之间的信号通常具有挑战性。我们提出了一种方法,我们称之为PBS或信号概率,以简化和简化识别ChIP-seq数据集中可以找到真实信号的过程,无论信号是宽的还是窄的。我们展示了我们的方法如何用于可视化,量化和比较ChIP-seq数据集之间的信号,并轻松地将ChIP-seq数据与其他数据类型集成。我们展示了PBS如何既能独立存在,又能为其他常用的分析方法提供支持。我们预计,它的多功能性和直接的解释将证明在许多应用中是有用的。
Histone ChIP-seq is one of the primary methods for charting the cellular epigenomic landscape, the components of which play a critical regulatory role in gene expression. Analyzing the activity of regulatory elements across datasets and cell types can be challenging due to shifting peak positions and normalization artifacts resulting from, for example, differing read depths, ChIP efficiencies, and target sizes. Moreover, broad regions of enrichment seen in repressive histone marks often evade detection by commonly used peak callers. Here, we present a simple and versatile method for identifying enriched regions in ChIP-seq data that relies on estimating a gamma distribution fit to non-overlapping 5kB genomic bins to establish a global background. We use this distribution to assign a probability of being signal (PBS) between zero and one to each 5 kB bin. This approach, while lower in resolution than typical peak-calling methods, provides a straightforward way to identify enriched regions and compare enrichments among multiple datasets, by transforming the data to values that are universally normalized and can be readily visualized and integrated with downstream analysis methods. We demonstrate applications of PBS for both broad and narrow histone marks, and provide several illustrations of biological insights which can be gleaned by integrating PBS scores with downstream data types. Histone modifications are key to epigenetic regulation of gene expression. Their genomic distributions are most commonly measured using ChIP-seq, a next-generation sequencing method which results in pileups of sequencing reads near regions with modified histones. Comparing the signal between histone modification profiles in different cellular contexts is often challenging due to the lack of a clear standard for interpreting signal strength, the inherent variability in the positions of histones, and the difficulty of detecting very broad regions that are only modestly enriched compared to the background. We present a method, which we call PBS, or probability of being signal, to simplify and streamline the process of identifying where true signal can be found in ChIP-seq datasets, regardless of whether the signal is broad or narrow. We demonstrate how our method can be used to visualize, quantify and compare signal among ChIP-seq datasets, and to easily integrate ChIP-seq data with additional data types. We show how PBS can both stand on its own, and also add power to other commonly used analysis methods. We anticipate that its versatility and straightforward interpretation will prove useful in many applications.
鉴定雌激素受体结合与乳腺癌的临床结局有关。
DOI: 10.1038/nature10730
发表时间: 2012-01-04
期刊: NATURE
影响因子: 64.8
作者:
Ross-Innes, Caryn S.;Stark, Rory;Teschendorff, Andrew E.;Holmes, Kelly A.;Ali, H. Raza;Dunning, Mark J.;Brown, Gordon D.;Gojis, Ondrej;Ellis, Ian O.;Green, Andrew R.;Ali, Simak;Chin, Suet-Feung;Palmieri, Carlo;Caldas, Carlos;Carroll, Jason S.
通讯作者: Carroll, Jason S.
DOI: 10.1016/j.cell.2020.07.030
发表时间: 2020-09-17
期刊: Cell
影响因子: 64.5
作者:
Johnstone SE;Reyes A;Qi Y;Adriaens C;Hegazi E;Pelka K;Chen JH;Zou LS;Drier Y;Hecht V;Shoresh N;Selig MK;Lareau CA;Iyer S;Nguyen SC;Joyce EF;Hacohen N;Irizarry RA;Zhang B;Aryee MJ;Bernstein BE
通讯作者: Bernstein BE
DOI: 10.4161/onci.19964
发表时间: 2012-07-01
期刊: Oncoimmunology
影响因子: 7.2
作者:
Kohanbash G;Ishikawa E;Fujita M;Ikeura M;McKaveney K;Zhu J;Sakaki M;Sarkar SN;Okada H
通讯作者: Okada H
DOI: 10.1371/journal.pcbi.1003326
发表时间: 2013
影响因子: 4.3
作者:
Bailey T;Krajewski P;Ladunga I;Lefebvre C;Li Q;Liu T;Madrigal P;Taslim C;Zhang J
通讯作者: Zhang J
DOI: 10.1038/ng.3404
发表时间: 2015-11
期刊: Nature genetics
影响因子: 30.8
作者:
Finucane HK;Bulik-Sullivan B;Gusev A;Trynka G;Reshef Y;Loh PR;Anttila V;Xu H;Zang C;Farh K;Ripke S;Day FR;ReproGen Consortium;Schizophrenia Working Group of the Psychiatric Genomics Consortium;RACI Consortium;Purcell S;Stahl E;Lindstrom S;Perry JR;Okada Y;Raychaudhuri S;Daly MJ;Patterson N;Neale BM;Price AL
通讯作者: Price AL