Bayesian hierarchical clustering for microarray time series data with replicates and outlier measurements.

Bayesian hierarchical clustering for microarray time series data with replicates and outlier measurements.
复制标题

DOI:
10.1186/1471-2105-12-399
复制
发表时间:
2011-10-13
期刊:
影响因子:
3
通讯作者:
Wild DL
Wild DL
中科院分区:
生物学4区
文献类型:
--
作者:
Cooke EJ;Savage RS;Kirk PD;Darkins R;Wild DL

文献摘要

参考文献

被引文献

相似文献

后基因组分子生物学导致了数据的爆炸,提供了大量基因,蛋白质和代谢物的测量。时间序列实验已经变得越来越普遍,需要开发新的分析工具来捕获所产生的数据结构。在一个或多个时间点的离群值测量提出了一个重大的挑战,而潜在的有价值的复制信息往往被忽略现有的技术。我们提出了一个基于生成模型的贝叶斯层次聚类算法的微阵列时间序列,采用高斯过程回归捕捉数据的结构。通过使用混合模型的可能性,我们的方法允许一小部分的数据被建模为离群值测量,并采用经验贝叶斯方法,使用重复的观察通知先验分布的噪声方差。该方法自动学习最佳的聚类数,并可以将非均匀采样的时间点。使用各种各样的实验数据集,我们表明,我们的算法始终产生更高的质量和更有生物意义的集群比目前最先进的方法。我们强调了建模离群值的重要性,通过证明嘈杂的基因可以与其他类似的生物功能的基因分组。我们证明了包括复制信息的重要性,我们发现这使得额外的不同的表达谱的歧视。通过将离群值测量和重复值,这种聚类算法的时间序列微阵列数据提供了一个步骤,更好地处理固有的噪声测量高通量基因组技术。Timeseries BHC作为R包'BHC'(版本1.5)的一部分提供,可通过http://www.bioconductor.org/packages/release/bioc/html/BHC.html?从Bioconductor(版本2.9及以上)下载pagewanted=全部。
Post-genomic molecular biology has resulted in an explosion of data, providing measurements for large numbers of genes, proteins and metabolites. Time series experiments have become increasingly common, necessitating the development of novel analysis tools that capture the resulting data structure. Outlier measurements at one or more time points present a significant challenge, while potentially valuable replicate information is often ignored by existing techniques. We present a generative model-based Bayesian hierarchical clustering algorithm for microarray time series that employs Gaussian process regression to capture the structure of the data. By using a mixture model likelihood, our method permits a small proportion of the data to be modelled as outlier measurements, and adopts an empirical Bayes approach which uses replicate observations to inform a prior distribution of the noise variance. The method automatically learns the optimum number of clusters and can incorporate non-uniformly sampled time points. Using a wide variety of experimental data sets, we show that our algorithm consistently yields higher quality and more biologically meaningful clusters than current state-of-the-art methodologies. We highlight the importance of modelling outlier values by demonstrating that noisy genes can be grouped with other genes of similar biological function. We demonstrate the importance of including replicate information, which we find enables the discrimination of additional distinct expression profiles. By incorporating outlier measurements and replicate values, this clustering algorithm for time series microarray data provides a step towards a better treatment of the noise inherent in measurements from high-throughput genomic technologies. Timeseries BHC is available as part of the R package 'BHC' (version 1.5), which is available for download from Bioconductor (version 2.9 and above) via http://www.bioconductor.org/packages/release/bioc/html/BHC.html?pagewanted=all.
DOI: 10.1016/s1097-2765(00)80114-8
发表时间: 1998-07-01
期刊: MOLECULAR CELL
影响因子: 16
作者:
Cho, RJ;Campbell, MJ;Davis, RW
通讯作者: Davis, RW
DOI: 10.1198/016214505000000187
发表时间: 2006-03-01
影响因子: 3.7
作者:
Heard, NA;Holmes, CC;Stephens, DA
通讯作者: Stephens, DA
DOI: 10.1093/bioinformatics/btq022
发表时间: 2010-03-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Liu, Qiang;Lin, Kevin K.;Ihler, Alexander
通讯作者: Ihler, Alexander
DOI: 10.1038/nature06955
发表时间: 2008-06-12
期刊: NATURE
影响因子: 64.8
作者:
Orlando, David A.;Lin, Charles Y.;Haase, Steven B.
通讯作者: Haase, Steven B.
DOI: 10.1101/gad.1450606
发表时间: 2006-08-15
影响因子: 10.5
作者:
Pramila, Tata;Wu, Wei;Breeden, Linda L.
通讯作者: Breeden, Linda L.