Deconvolving sequence features that discriminate between overlapping regulatory annotations.

Deconvolving sequence features that discriminate between overlapping regulatory annotations.
复制标题

DOI:
10.1371/journal.pcbi.1005795
复制
发表时间:
2017-10
影响因子:
4.3
通讯作者:
Mahony S
Mahony S
中科院分区:
生物学2区
文献类型:
--
作者:
Kakumanu A;Velasco S;Mazzoni E;Mahony S

文献摘要

参考文献

被引文献

相似文献

具有调控潜力的基因组基因座可以用各种性质注释。例如,由给定的转录因子(TF)结合的基因组位点可以根据它们是在已知启动子的近端还是远端来划分。可以根据细胞类型和它们活跃的条件进一步标记位点。给定这样的标记位点的集合,自然会询问哪些序列特征与每个注释标签相关联。然而,发现这样的标记特异性序列特征经常被标记之间的重叠混淆;例如,如果对给定细胞类型特异性的调控位点也更可能是启动子近端的,则难以评估在该组位点中鉴定的基序是与细胞类型相关还是与启动子相关。为了应对这一挑战,我们开发了SeqUnwinder,这是一种原理性方法,用于解卷积与重叠注释标签相关的可解释的判别序列特征。我们使用三个例子来展示SeqUnwinder的新分析能力。首先,SeqUnwinder能够从与初始胚胎干细胞中的染色质状态相关的特征中解开与运动神经元编程期间TF的动态结合行为相关的序列特征。其次,我们的特点不同的序列特性的多条件和细胞特异性TF结合位点后,控制不均匀协会与启动子接近。最后,我们展示了SeqUnwinder的可扩展性,以发现来自一个或多个ENCODE细胞系中显示DNase I超敏性的十万多个基因组位点的细胞特异性序列特征。转录因子蛋白通过识别基因组调控区中的短DNA序列模式并与之相互作用来控制基因表达。目前的基因组学实验使我们能够在整个基因组中找到与特定生化活性相关的调控区域;例如,特定转录因子与给定细胞类型中的基因组相互作用的所有区域。给定一组调控区,我们的目标通常是发现在该组中比在其他区域中更常见的短DNA序列模式。进行这样的“DNA基序发现”分析可以给我们提示,在所分析的细胞类型中决定基因调控的模式。在这里,我们描述了一种新的方法,DNA模体发现称为SeqUnwinder。我们的方法分析了调控区的集合,每个调控区都根据不同的生物学特性进行了标记。例如,标记可以对应于其中调节区是活性的各种细胞类型。然后,SeqUnwinder进行机器学习分析,以解开每个标签的DNA序列特征(例如,将每个细胞类型中的调控区与其他细胞类型区分开的特征)。SeqUnwinder是第一种能够分析包含多个重叠标签的监管区域集合的方法。
Genomic loci with regulatory potential can be annotated with various properties. For example, genomic sites bound by a given transcription factor (TF) can be divided according to whether they are proximal or distal to known promoters. Sites can be further labeled according to the cell types and conditions in which they are active. Given such a collection of labeled sites, it is natural to ask what sequence features are associated with each annotation label. However, discovering such label-specific sequence features is often confounded by overlaps between the labels; e.g. if regulatory sites specific to a given cell type are also more likely to be promoter-proximal, it is difficult to assess whether motifs identified in that set of sites are associated with the cell type or associated with promoters. In order to meet this challenge, we developed SeqUnwinder, a principled approach to deconvolving interpretable discriminative sequence features associated with overlapping annotation labels. We demonstrate the novel analysis abilities of SeqUnwinder using three examples. Firstly, SeqUnwinder is able to unravel sequence features associated with the dynamic binding behavior of TFs during motor neuron programming from features associated with chromatin state in the initial embryonic stem cells. Secondly, we characterize distinct sequence properties of multi-condition and cell-specific TF binding sites after controlling for uneven associations with promoter proximity. Finally, we demonstrate the scalability of SeqUnwinder to discover cell-specific sequence features from over one hundred thousand genomic loci that display DNase I hypersensitivity in one or more ENCODE cell lines. Transcription factor proteins control gene expression by recognizing and interacting with short DNA sequence patterns in regulatory regions on the genome. Current genomics experiments allow us to find regulatory regions associated with a particular biochemical activity over the entire genome; for example, all regions where a particular transcription factor interacts with the genome in a given cell type. Given a collection of regulatory regions, we often aim to discover short DNA sequence patterns that are more common in the collection than in other regions. Performing such “DNA motif-finding” analysis can give us hints about the patterns that determine gene regulation in the analyzed cell type. Here we describe a new method for DNA motif-finding called SeqUnwinder. Our approach analyzes collections of regulatory regions where each has been labeled according to various biological properties. For example, the labels could correspond to various cell types in which the regulatory region is active. SeqUnwinder then performs machine-learning analysis to unravel DNA sequence features that are characteristic of each label (e.g. features that distinguish regulatory regions in each cell type from other cell types). SeqUnwinder is the first method to enable analysis of regulatory region collections that contain several overlapping labels.
DOI: 10.1093/nar/gkt1249
发表时间: 2014-03
影响因子: 14.9
作者:
Kheradpour P;Kellis M
通讯作者: Kellis M
DOI: 10.1016/j.celrep.2014.08.046
发表时间: 2014-10-09
期刊: Cell reports
影响因子: 8.8
作者:
Alder O;Cullum R;Lee S;Kan AC;Wei W;Yi Y;Garside VC;Bilenky M;Griffith M;Morrissy AS;Robertson GA;Thiessen N;Zhao Y;Chen Q;Pan D;Jones SJM;Marra MA;Hoodless PA
通讯作者: Hoodless PA
DOI: 10.1016/j.cell.2006.12.048
发表时间: 2007-03-23
期刊: CELL
影响因子: 64.5
作者:
Kim, Tae Hoon;Abdullaev, Ziedulla K.;Ren, Bing
通讯作者: Ren, Bing
DOI: 10.1101/gr.082800.108
发表时间: 2009-01-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Cuddapah, Suresh;Jothi, Raja;Zhao, Keji
通讯作者: Zhao, Keji
DOI: 10.1016/j.cell.2012.01.030
发表时间: 2012-02-03
期刊: CELL
影响因子: 64.5
作者:
Junion, Guillaume;Spivakov, Mikhail;Furlong, Eileen E. M.
通讯作者: Furlong, Eileen E. M.