Penalized and weighted K-means for clustering with scattered objects and prior information in high-throughput biological data

Penalized and weighted K-means for clustering with scattered objects and prior information in high-throughput biological data
复制标题

DOI:
10.1093/bioinformatics/btm320
复制
发表时间:
2007-09-01
期刊:
影响因子:
5.8
通讯作者:
Tseng, George C.
Tseng, George C.
中科院分区:
生物学3区
文献类型:
--
作者:
Tseng, George C.

文献摘要

被引文献

相似文献

动机:聚类分析是研究高通量生物数据的最重要的数据挖掘工具之一。传统的聚类算法在高维复杂的情况下,由于存在大量不应该被聚类的分散对象,从而影响了算法的性能。通常,来自数据库或先前实验的额外先验知识也可用于分析。排除分散的对象和合并现有的先验信息是可取的,以提高聚类性能。结果:在本文中,提出了一类损失函数的聚类分析,并应用于高通量的基因组和蛋白质组数据。K-means的两个主要扩展涉及:惩罚和加权。附加惩罚项用于允许一组分散的对象不被聚类。引入权重来说明要识别的优选或禁止的聚类模式的先验信息。探讨了它们与高斯混合模型分类似然性的关系。结合良好的先验信息也被证明可以改善聚类中的全局优化问题。应用所提出的方法模拟数据,以及高通量数据集的串联质谱(MS/MS)和微阵列实验。我们的研究结果表明,它的上级性能优于大多数现有的方法和它的计算简单性和可扩展性,在大型复杂的生物数据集的应用。
Motivation: Cluster analysis is one of the most important data mining tools for investigating high-throughput biological data. The existence of many scattered objects that should not be clustered has been found to hinder performance of most traditional clustering algorithms in such a high- dimensional complex situation. Very often, additional prior knowledge from databases or previous experiments is also available in the analysis. Excluding scattered objects and incorporating existing prior information are desirable to enhance the clustering performance.Results: In this article, a class of loss functions is proposed for cluster analysis and applied in high- throughput genomic and proteomic data. Two major extensions from K-means are involved: penalization and weighting. The additive penalty term is used to allow a set of scattered objects without being clustered. Weights are introduced to account for prior information of preferred or prohibited cluster patterns to be identified. Their relationship with the classification likelihood of Gaussian mixture models is explored. Incorporation of good prior information is also shown to improve the global optimization issue in clustering. Applications of the proposed method on simulated data as well as high- throughput data sets from tandem mass spectrometry (MS/MS) and microarray experiments are presented. Our results demonstrate its superior performance over most existing methods and its computational simplicity and extensibility in the application of large complex biological data sets.