Cluster Analysis and Related Issues

Cluster Analysis and Related Issues
复制标题

DOI:
10.1142/9789814343138_0001
复制
发表时间:
1993-12
期刊:
--
影响因子:
--
通讯作者:
R. Dubes
R. Dubes
中科院分区:
其他
文献类型:
--
作者:
R. Dubes

文献摘要

被引文献

相似文献

本章解释聚类分析如何在计算机视觉和模式识别等应用中组织信息。信息表示为多维特征空间中的点,其中每个坐标表示一个度量。探索性数据分析的一些工具进行了讨论,重点是来自协方差矩阵的线性预测。两种类型的集群进行审查-层次和分区。层次聚类导致数据的嵌套分区。SAHN算法的层次聚类的定义和一些共同的特点进行了解释。分区聚类将数据排列在单独的聚类中,与K-Means算法一样。本章最后讨论了验证,集中在外部和内部的有效性测试和测试的集群数量。为进一步阅读提供了参考书目。
This chapter explains how cluster analysis organizes information in applications such as Computer Vision and Pattern Recognition. Information is represented as points in multidimensional feature spaces where each coordinate represents a measurement. Some tools from exploratory data analysis are discussed, with an emphasis on linear projections derived from the covariance matrix. Two types of clustering are reviewed — hierarchical and partitional. Hierarchical clustering leads to nested partitions of the data. SAHN algorithms for hierarchical clustering are defined and some of the common characteristics are explained. Partitional clustering arranges data in separate clusters, as with the K-Means algorithm. The chapter ends with a discussion of validation that centers on external and internal tests of validity and tests for the number of clusters. A bibliography is provided for further reading.