NETWORK MODELLING OF TOPOLOGICAL DOMAINS USING HI-C DATA.

NETWORK MODELLING OF TOPOLOGICAL DOMAINS USING HI-C DATA.
复制标题

DOI:
10.1214/19-aoas1244
复制
发表时间:
2019-09
期刊:
The annals of applied statistics
影响因子:
--
通讯作者:
Bickel PJ
Bickel PJ
中科院分区:
其他
文献类型:
--
作者:
Wang YXR;Sarkar P;Ursu O;Kundaje A;Bickel PJ

文献摘要

参考文献

被引文献

相似文献

染色体构象捕捉实验如Hi-C用于绘制基因组的三维空间组织图。3D组织的一个特定特征被称为拓扑关联结构域(TADS),它是紧密相互作用的、连续的染色质区域,在调节基因表达方面发挥着重要作用。目前已经提出了几种检测TADS的算法。特别是,Hi-C数据的结构自然启发了社区检测方法的应用。然而,群落发现的一个缺点是,大多数方法认为网络中节点的可交换性是理所当然的,而这种情况下的节点,即染色体上的位置,是不可交换的。我们提出了一种利用Hi-C数据检测TADS的网络模型,该模型考虑了这种不可交换性。此外,我们的模型明确地利用细胞类型特定的CTCF结合位点作为生物协变量,并可用于识别多种细胞类型的保守TADS。该模型产生了一个似然目标,可以通过松弛有效地进行优化。我们还证明了,当适当地初始化时,该模型以很高的概率找到潜在的TAD结构。使用模拟数据,我们展示了我们的方法的优势和流行的社区检测方法,如谱聚类,在这一应用中的注意事项。将我们的方法应用于真实的Hi-C数据,我们证明了所识别的区域具有理想的表观遗传学特征,并对不同类型的细胞进行了比较。
Chromosome conformation capture experiments such as Hi-C are used to map the three-dimensional spatial organization of genomes. One specific feature of the 3D organization is known as topologically associating domains (TADs), which are densely interacting, contiguous chromatin regions playing important roles in regulating gene expression. A few algorithms have been proposed to detect TADs. In particular, the structure of Hi-C data naturally inspires application of community detection methods. However, one of the drawbacks of community detection is that most methods take exchangeability of the nodes in the network for granted; whereas the nodes in this case, that is, the positions on the chromosomes, are not exchangeable. We propose a network model for detecting TADs using Hi-C data that takes into account this nonexchangeability. in addition, our model explicitly makes use of cell-type specific CTCF binding sites as biological covariates and can be used to identify conserved TADs across multiple cell types. The model leads to a likelihood objective that can be efficiently optimized via relaxation. We also prove that when suitably initialized, this model finds the underlying TAD structure with high probability. using simulated data, we show the advantages of our method and the caveats of popular community detection methods, such as spectral clustering, in this application. Applying our method to real Hi-C data, we demonstrate the domains identified have desirable epigenetic features and compare them across different cell types.
用于分析HI-C数据的二维分割。
DOI: 10.1093/bioinformatics/btu443
发表时间: 2014-09-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lévy-Leduc C;Delattre M;Mary-Huard T;Robin S
通讯作者: Robin S
DOI: 10.1038/nmeth.4560
发表时间: 2018-03
期刊: Nature methods
影响因子: 48
作者:
Norton HK;Emerson DJ;Huang H;Kim J;Titus KR;Gu S;Bassett DS;Phillips-Cremins JE
通讯作者: Phillips-Cremins JE
DOI: 10.1186/1748-7188-9-14
发表时间: 2014
期刊: Algorithms for molecular biology : AMB
影响因子: --
作者:
Filippova D;Patro R;Duggal G;Kingsford C
通讯作者: Kingsford C
DOI: 10.1083/jcb.200909127
发表时间: 2009-12-14
影响因子: 7.8
作者:
Meaburn, Karen J.;Gudla, Prabhakar R.;Misteli, Tom
通讯作者: Misteli, Tom
DOI: 10.1016/j.molcel.2012.08.031
发表时间: 2012-11-09
期刊: MOLECULAR CELL
影响因子: 16
作者:
Hou, Chunhui;Li, Li;Qin, Zhaohui S.;Corces, Victor G.
通讯作者: Corces, Victor G.