Identifying topologically associating domains and subdomains by Gaussian Mixture model And Proportion test.

Identifying topologically associating domains and subdomains by Gaussian Mixture model And Proportion test.
复制标题

DOI:
10.1038/s41467-017-00478-8
复制
发表时间:
2017-09-14
影响因子:
16.6
通讯作者:
Tan K
Tan K
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Yu W;He B;Tan K

文献摘要

参考文献

相似文献

基因组的空间组织在基因表达调控中起着关键作用。最近的染色质相互作用作图研究表明,拓扑相关结构域和子结构域是三维基因组的基本组成部分。识别这种层次结构是理解基因组三维结构-功能关系的关键一步。现有的计算算法缺乏对区域预测的统计评估,并且对于高分辨率的Hi-C数据计算效率低下。我们引入高斯混合模型和比例测试(GMAP)算法来解决上述挑战。使用模拟和实验的Hi-C数据,我们表明GMAP识别的域比三种最先进的方法更符合多条支持证据。将GMAP应用于正常细胞和癌细胞,揭示了亚结构域边界与结构域边界相比的几个独特特征,包括其跨细胞类型的更高动力学和癌症体细胞突变的富集。基因组的空间组织在基因表达调控中起着至关重要的作用。本文引入GMAP,即高斯混合模型和比例检验来识别Hi-C数据中的拓扑关联域和子域。
The spatial organization of the genome plays a critical role in regulating gene expression. Recent chromatin interaction mapping studies have revealed that topologically associating domains and subdomains are fundamental building blocks of the three-dimensional genome. Identifying such hierarchical structures is a critical step toward understanding the three-dimensional structure–function relationship of the genome. Existing computational algorithms lack statistical assessment of domain predictions and are computationally inefficient for high-resolution Hi-C data. We introduce the Gaussian Mixture model And Proportion test (GMAP) algorithm to address the above-mentioned challenges. Using simulated and experimental Hi-C data, we show that domains identified by GMAP are more consistent with multiple lines of supporting evidence than three state-of-the-art methods. Application of GMAP to normal and cancer cells reveals several unique features of subdomain boundary as compared to domain boundary, including its higher dynamics across cell types and enrichment for somatic mutations in cancer. Spatial organization of the genome plays a crucial role in regulating gene expression. Here the authors introduce GMAP, the Gaussian Mixture model And Proportion test, to identify topologically associating domains and subdomains in Hi-C data.
DOI: 10.1038/nbt.1621
发表时间: 2010-05
影响因子: 46.9
作者:
Trapnell C;Williams BA;Pertea G;Mortazavi A;Kwan G;van Baren MJ;Salzberg SL;Wold BJ;Pachter L
通讯作者: Pachter L
DOI: 10.1016/j.cell.2013.04.053
发表时间: 2013-06-06
期刊: Cell
影响因子: 64.5
作者:
Phillips-Cremins JE;Sauria ME;Sanyal A;Gerasimova TI;Lajoie BR;Bell JS;Ong CT;Hookway TA;Guo C;Sun Y;Bland MJ;Wagstaff W;Dalton S;McDevitt TC;Sen R;Dekker J;Taylor J;Corces VG
通讯作者: Corces VG
用于分析HI-C数据的二维分割。
DOI: 10.1093/bioinformatics/btu443
发表时间: 2014-09-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lévy-Leduc C;Delattre M;Mary-Huard T;Robin S
通讯作者: Robin S
DOI: 10.1186/1748-7188-9-14
发表时间: 2014
期刊: Algorithms for molecular biology : AMB
影响因子: --
作者:
Filippova D;Patro R;Duggal G;Kingsford C
通讯作者: Kingsford C
DOI: 10.1073/pnas.1320308111
发表时间: 2014-05-27
影响因子: 11.1
作者:
He, Bing;Chen, Changya;Tan, Kai
通讯作者: Tan, Kai