Analyzing gene expression time-courses

Analyzing gene expression time-courses
复制标题

DOI:
10.1109/tcbb.2005.31
复制
发表时间:
2005-07-01
影响因子:
4.5
通讯作者:
Schönhuth, A
Schönhuth, A
中科院分区:
工程技术3区
文献类型:
--
作者:
Schliep, A;Costa, IG;Schönhuth, A

文献摘要

被引文献

相似文献

随着时间的推移测量基因表达可以为基本细胞过程提供重要的见解。识别具有相似表达时间进程的基因组是分析中至关重要的第一步。由于基因在这些细胞过程中具有多个不同的作用,因此生物学相关的群体经常重叠,这对于经典聚类方法来说是一个难题。我们使用混合模型来规避这个主要问题,并使用隐马尔可夫模型(HMM)作为有效且灵活的组件。我们表明,通过修改期望最大化(EM)算法,可以使用额外的标记数据(对混合物的部分监督学习)来解决随后的估计问题。通过对贝叶斯模型合并的修改获得了混合估计的良好起点,这使我们能够学习初始 HMM 的集合。我们使用简单的信息论解码启发式从混合物中推断出组,该解码启发式量化了组分配中的模糊程度。高质量的注释数据显示了其有效性。由于我们提出的 HMM 通过设计捕获异步行为,因此我们找到的组也是异步的。同步子群是从基于维特比路径的新算法获得的,我们通过与以前的方法进行有利的比较,展示了我们的 HMM 混合方法对生物和模拟数据的适用性。实现该方法的软件可以根据 GPL 从 http://ghmm.org/gql 免费获得。
Measuring gene expression over time can provide important insights into basic cellular processes. Identifying groups of genes with similar expression time-courses is a crucial first step in the analysis. As biologically relevant groups frequently overlap, due to genes having several distinct roles in those cellular processes, this is a difficult problem for classical clustering methods. We use a mixture model to circumvent this principal problem, with hidden Markov models (HMMs) as effective and flexible components. We show that the ensuing estimation problem can be addressed with additional labeled data-partially supervised learning of mixtures-through a modification of the Expectation-Maximization (EM) algorithm. Good starting points for the mixture estimation are obtained through a modification to Bayesian model merging, which allows us to learn a collection of initial HMMs. We infer groups from mixtures with a simple information-theoretic decoding heuristic, which quantifies the level of ambiguity in group assignment. The effectiveness is shown with high-quality annotation data. As the HMMs we propose capture asynchronous behavior by design, the groups we find are also asynchronous. Synchronous subgroups are obtained from a novel algorithm based on Viterbi paths, We show the suitability of our HMM mixture approach on biological and simulated data and through the favorable comparison with previous approaches. A software implementing the method is freely available under the GPL from http://ghmm.org/gql.