Predicting human nucleosome occupancy from primary sequence.

Predicting human nucleosome occupancy from primary sequence.
复制标题

DOI:
10.1371/journal.pcbi.1000134
复制
发表时间:
2008-08-22
影响因子:
4.3
通讯作者:
Noble WS
Noble WS
中科院分区:
生物学2区
文献类型:
--
作者:
Gupta S;Dennis J;Thurman RE;Kingston R;Stamatoyannopoulos JA;Noble WS

文献摘要

参考文献

被引文献

相似文献

核小体是染色质的基本重复单位,并且包括活的真核基因组的结构构建块。微球菌核酸酶(MNase)长期以来一直被用来描绘核小体的组织。酵母染色质中基于微阵列的核小体作图实验揭示了核小体的规则间隔的翻译定相。这些数据已被用于训练序列定向核小体定位的计算模型,其已经识别出普遍存在的强内在核小体定位信号。在这里,我们成功地将这种方法应用于人类染色质的核小体定位实验。由人类训练的和酵母训练的模型做出的预测是强相关的,这表明基于序列的核小体占据的确定具有共同的机制。此外,我们还观察到,在弱消化与重消化MNase样品的实验数据上训练的分类器之间存在显著的互补性。在前一种情况下,所得到的模型准确地识别核小体形成序列;在后一种情况下,分类器在识别无核小体区域方面表现出色。使用这个模型,我们能够确定核小体形成和核小体不利序列的几个特征。首先,通过将来自从头应用于人ENCODE区域的每个分类器的结果组合,分类器揭示了核小体形成和核小体不利序列的不同序列组成和周期性特征。短的二核苷酸重复序列作为核小体不利序列的标志出现,而核小体形成序列包含短的GC碱基对周期性运行。第二,我们表明,核小体定相是最常见的预测侧翼无核小体区域。结果表明,核小体在体内定位的主要机制是边界事件驱动的,并肯定了经典的核小体组织的统计定位理论。在细胞核内,DNA被包裹成一个复杂的分子结构,称为染色质,其基本单位是150 bp的DNA,组织在8个组蛋白蛋白复合物周围,称为核小体。了解核小体的局部组织对于了解染色质如何影响基因调控至关重要。在这里,我们描述了一个计算模型,预测从DNA序列的核小体的位置。我们使用来自人类细胞系的数据训练模型,并将该模型系统地应用于人类基因组的1%。我们发现,先前描述的从酵母数据训练的模型与人类训练的模型密切相关,这表明了基于序列的核小体占有率测定的共同机制。此外,我们观察到使用弱消化和强消化样本的数据训练的模型之间的惊人互补性:一种类型的模型识别无核小体区域,而另一种识别定位良好的核小体。最后,我们预测的核小体在人类基因组中的位置的分析,使我们能够识别核小体形成和抑制序列的共同特征。总的来说,我们的结果是一致的核小体组织的经典统计定位理论。
Nucleosomes are the fundamental repeating unit of chromatin and comprise the structural building blocks of the living eukaryotic genome. Micrococcal nuclease (MNase) has long been used to delineate nucleosomal organization. Microarray-based nucleosome mapping experiments in yeast chromatin have revealed regularly-spaced translational phasing of nucleosomes. These data have been used to train computational models of sequence-directed nuclesosome positioning, which have identified ubiquitous strong intrinsic nucleosome positioning signals. Here, we successfully apply this approach to nucleosome positioning experiments from human chromatin. The predictions made by the human-trained and yeast-trained models are strongly correlated, suggesting a shared mechanism for sequence-based determination of nucleosome occupancy. In addition, we observed striking complementarity between classifiers trained on experimental data from weakly versus heavily digested MNase samples. In the former case, the resulting model accurately identifies nucleosome-forming sequences; in the latter, the classifier excels at identifying nucleosome-free regions. Using this model we are able to identify several characteristics of nucleosome-forming and nucleosome-disfavoring sequences. First, by combining results from each classifier applied de novo across the human ENCODE regions, the classifier reveals distinct sequence composition and periodicity features of nucleosome-forming and nucleosome-disfavoring sequences. Short runs of dinucleotide repeat appear as a hallmark of nucleosome-disfavoring sequences, while nucleosome-forming sequences contain short periodic runs of GC base pairs. Second, we show that nucleosome phasing is most frequently predicted flanking nucleosome-free regions. The results suggest that the major mechanism of nucleosome positioning in vivo is boundary-event-driven and affirm the classical statistical positioning theory of nucleosome organization. Inside the nucleus, DNA is wrapped into a complex molecular structure called chromatin, whose fundamental unit is ∼150 bp of DNA organized around the eight-histone protein complex known as the nucleosome. Understanding the local organization of nucleosomes is critical for understanding how chromatin impacts gene regulation. Here, we describe a computational model that predicts nucleosome placement from DNA sequence. We train the model using data derived from human cell lines, and we apply the model systematically to 1% of the human genome. We show that previously described models trained from yeast data correlate strongly with the human-trained model, suggesting a common mechanism for sequence-based determination of nucleosome occupancy. In addition, we observe a striking complementarity between models trained using data from weakly and strongly digested samples: one type of model recognizes nucleosome-free regions, whereas the other identifies well-positioned nucleosomes. Finally, our analysis of predicted nucleosome positions in the human genome allows us to identify common features of nucleosome-forming and inhibitory sequences. Overall, our results are consistent with the classical statistical positioning theory of nucleosome organization.
DOI: 10.1016/0092-8674(88)90466-7
发表时间: 1988-02-26
期刊: CELL
影响因子: 64.5
作者:
HSIEH, CH;GRIFFITH, JD
通讯作者: GRIFFITH, JD
DOI: 10.1038/329263a0
发表时间: 1987-09-17
期刊: NATURE
影响因子: 64.8
作者:
HOGAN, ME;AUSTIN, RH
通讯作者: AUSTIN, RH
DOI: 10.1016/j.gde.2004.01.007
发表时间: 2004-04-01
影响因子: 4
作者:
Flaus, A;Owen-Hughes, T
通讯作者: Owen-Hughes, T
DOI: 10.1002/j.1460-2075.1986.tb04551.x
发表时间: 1986-10-01
期刊: EMBO JOURNAL
影响因子: 11.4
作者:
ALMER, A;HORZ, W
通讯作者: HORZ, W
DOI: 10.1016/0022-2836(85)90396-1
发表时间: 1985-12-20
影响因子: 5.6
作者:
DREW, HR;TRAVERS, AA
通讯作者: TRAVERS, AA