Learning chromatin states with factorized information criteria

Learning chromatin states with factorized information criteria
复制标题

DOI:
10.1093/bioinformatics/btv163
复制
发表时间:
2015-08-01
期刊:
影响因子:
5.8
通讯作者:
Asai, Kiyoshi
Asai, Kiyoshi
中科院分区:
生物学3区
文献类型:
--
作者:
Hamada, Michiaki;Ono, Yukiteru;Asai, Kiyoshi

文献摘要

被引文献

相似文献

动机:最近的研究表明,基因组和具有表观遗传修饰的基因组,即所谓的表观基因组,在各种生物功能中发挥重要作用,如转录和DNA复制,修复和重组。众所周知,核小体的组蛋白修饰(例如甲基化和乙酰化)的特定组合诱导对应于染色质的特定功能的染色质状态。虽然下一代测序(NGS)技术的出现,使整个基因组的表观遗传信息的测量在高分辨率,染色质状态的变化还没有完全characterized.Results:在这项研究中,我们提出了一种方法来估计的染色质状态表示的全基因组染色质标记确定的NGS技术。所提出的方法自动估计染色质状态的数量和特征的隐马尔可夫模型(HMM)的基础上,结合最近提出的模型选择技术,因式分解的信息标准的每个状态。由于该方法仅依赖于两个可调参数,并尽可能避免了启发式程序,因此有望提供无偏模型。模拟数据集的计算实验表明,我们的方法自动学习一个适当的模型,即使在依赖于贝叶斯信息标准的方法无法学习模型结构的情况下。此外,我们在三个真实的数据集上将我们的方法与ChromHMM进行了全面的比较,并表明我们的方法比ChromHMM对这些数据集估计了更多的染色质状态。
Motivation: Recent studies have suggested that both the genome and the genome with epigenetic modifications, the so-called epigenome, play important roles in various biological functions, such as transcription and DNA replication, repair, and recombination. It is well known that specific combinations of histone modifications (e.g. methylations and acetylations) of nucleosomes induce chromatin states that correspond to specific functions of chromatin. Although the advent of next-generation sequencing (NGS) technologies enables measurement of epigenetic information for entire genomes at high-resolution, the variety of chromatin states has not been completely characterized.Results: In this study, we propose a method to estimate the chromatin states indicated by genome-wide chromatin marks identified by NGS technologies. The proposed method automatically estimates the number of chromatin states and characterize each state on the basis of a hidden Markov model (HMM) in combination with a recently proposed model selection technique, factorized information criteria. The method is expected to provide an unbiased model because it relies on only two adjustable parameters and avoids heuristic procedures as much as possible. Computational experiments with simulated datasets show that our method automatically learns an appropriate model, even in cases where methods that rely on Bayesian information criteria fail to learn the model structures. In addition, we comprehensively compare our method to ChromHMM on three real datasets and show that our method estimates more chromatin states than ChromHMM for those datasets.