Robust inference of groups in gene expression time-courses using mixtures of HMMs

Robust inference of groups in gene expression time-courses using mixtures of HMMs
复制标题

DOI:
10.1093/bioinformatics/bth937
复制
发表时间:
2004-08-04
期刊:
影响因子:
5.8
通讯作者:
Schoenhuth, Alexander
Schoenhuth, Alexander
中科院分区:
生物学3区
文献类型:
--
作者:
Schliep, Alexander;Steinhoff, Christine;Schoenhuth, Alexander

文献摘要

被引文献

相似文献

动机:细胞过程的遗传调控经常使用大规模基因表达实验来研究,以观察表达随时间的变化。这种时态数据对经典的基于距离的聚类方法提出了挑战,因为它沿着时间轴具有水平依赖性。我们建议使用隐马尔可夫模型(HALGOT)来显式地对这些时间依赖性进行建模。在混合方法中使用的障碍,我们证明是上级聚类。此外,混合物是生物现实的更现实的模型,因为将基因明确划分为具有独特功能分配的簇是不可能的。混合物的使用增加了相对于噪声的鲁棒性,并允许在不同的分配模糊度水平下对组进行推断。一种简单的方法,部分监督学习,允许在训练期间受益于先前的生物学知识。我们的方法允许同时分析的循环和非循环基因和处理以及噪声和missing values.Results:我们证明了生物相关性检测的特定阶段的分组在HeLa的时间过程中的数据。使用模拟数据的基准,使用独立于我们的方法中的假设,显示出非常有利的结果相比,基线提供的k-means和两个先前的方法实现基于模型的聚类。研究结果强调了将先验知识,只要可用的好处。
Motivation: Genetic regulation of cellular processes is frequently investigated using large-scale gene expression experiments to observe changes in expression over time. This temporal data poses a challenge to classical distance-based clustering methods due to its horizontal dependencies along the time-axis. We propose to use hidden Markov models (HMMs) to explicitly model these time-dependencies. The HMMs are used in a mixture approach that we show to be superior over clustering. Furthermore, mixtures are a more realistic model of the biological reality, as an unambiguous partitioning of genes into clusters of unique functional assignment is impossible. Use of the mixture increases robustness with respect to noise and allows an inference of groups at varying level of assignment ambiguity. A simple approach, partially supervised learning, allows to benefit from prior biological knowledge during the training. Our method allows simultaneous analysis of cyclic and non-cyclic genes and copes well with noise and missing values.Results: We demonstrate biological relevance by detection of phase-specific groupings in HeLa time-course data. A benchmark using simulated data, derived using assumptions independent of those in our method, shows very favorable results compared to the baseline supplied by k-means and two prior approaches implementing model-based clustering. The results stress the benefits of incorporating prior knowledge, whenever available.