A program to perform Ward's clustering method on several regionalized variables

A program to perform Ward's clustering method on several regionalized variables
复制标题

DOI:
10.1016/j.cageo.2004.07.003
复制
发表时间:
2004-10-01
影响因子:
4.4
通讯作者:
Jarauta-Bragulat, E
Jarauta-Bragulat, E
中科院分区:
地球科学2区
文献类型:
--
作者:
Hervada-Sala, C;Jarauta-Bragulat, E

文献摘要

被引文献

相似文献

有许多统计技术可以找到数据和变量之间的相似或不同之处。聚类分析包含许多不同的技术,用于发现复杂数据集内的结构。聚类分析的目标是将数据或变量分组到聚类中,以便聚类中的元素之间具有高度的“自然关联”,而聚类之间则相互“相对不同”。为此,描述了许多标准:划分方法、任意原点方法、相互相似过程和层次聚类技术。Ward方法是应用最广泛的一种层次聚类法。地球科学研究一般涉及多变量和区域化的观测,这些观测可能是成分的,即百分比、浓度、mg/kg(Ppm)等数据。有时,了解这些数据是否必须划分为不同的子总体是一件有趣的事情。这个问题不能用传统的Ward方法来研究,因为样本不是独立的。在这种情况下,可以使用将Ward的聚类法扩展到空间相关样本。该方法基于广义马氏距离,该距离使用协方差和交叉协方差(或变异函数和交叉变异函数)矩阵。本文描述了对先前定义的这种方法的改进,其迭代和繁琐,因为在每一步都需要重新估计空间协方差结构。在本文中,我们保持在相同的理论框架内,但我们改进了方法,使用快速傅立叶变换方法来寻找协方差结构。由此,我们得到了改进型Ward聚类法对多个变量的推广。(C)2004爱思唯尔有限公司。保留所有权利。
There are many statistical techniques that allow finding similarities or differences among data and variables. Cluster analysis encompasses many diverse techniques for discovering structure within complex sets of data. The objective of cluster analysis is to group either the data or the variables into clusters such that the elements within a cluster have a high degree of "natural association" among themselves while clusters are "relatively distinct" from one another. To do so, many criteria have been described: partitioning methods, arbitrary origin methods, mutual similarity procedures and hierarchical clustering techniques. One of the most widespread hierarchical clustering methods is the Ward's method.Earth science studies deal in general with multivariate and regionalized observations which may be compositional, i.e. data such as percentages, concentrations, mg/kg (ppm). Sometimes, it is interesting to know whether these data have to be divided into different subpopulations. This problem cannot be studied with traditional Ward's method because samples are not independent. In that case, an extension of Ward's clustering method to spatially dependent samples can be used. This methodology is based on a generalized Mahalanobis distance, which uses the covariance and cross-covariance (or variogram and cross-variogram) matrices. This paper describes a refinement of this method previously defined, which was iterative and tedious, as it was necessary to re-estimate the spatial covariance structure at each step. In this paper, we stay within the same theoretical framework, but we improve the methodology using the fast fourier Transform method to find the covariance structure. Thus, we obtain a generalization to several variables of adapted Ward's clustering method. (C) 2004 Elsevier Ltd. All rights reserved.