Classification of geochemical data based on multivariate statistical analyses: Complementary roles of cluster, principal component, and independent component analyses

Classification of geochemical data based on multivariate statistical analyses: Complementary roles of cluster, principal component, and independent component analyses
复制标题

基于多元统计分析的地球化学数据分类:聚类分析、主成分分析和独立成分分析的互补作用

DOI:
10.1002/2016gc006663
复制
发表时间:
2017
期刊:
Geochem. Geophys. Geosyst.
影响因子:
--
通讯作者:
Kenta Ueki,
Kenta Ueki,
中科院分区:
--
文献类型:
--
作者:
Hikaru Iwamori;Kenta Yoshida;Hitomi Nakamura;Tatsu Kuwatani;Morihisa Hamada;Satoru Haraguchi;Kenta Ueki,

文献摘要

相似文献

识别地球化学问题中的数据结构,包括趋势和组/簇,对于从观测到的数据变异性讨论源和过程的起源至关重要。现代地球化学数据的数量和维数的不断增加,对多元统计分析方法提出了更高的要求。在本文中,我们展示了k均值聚类分析(KCA),主成分分析(PCA)和独立成分分析(伊卡)的关系和互补作用,以捕捉真实的数据结构。当通过初级标准化对数据进行预处理时(即,具有零平均值并通过标准偏差归一化),KCA和PCA提供基本上相同的结果,尽管前者返回离散空间中的解。当通过白化对数据进行预处理时(即,通过沿着主成分的特征值归一化),KCA和伊卡可以识别一组独立的趋势和组,而不管方差的幅度(功率)。作为一个例子,玄武岩同位素组成已被分析与KCA白化数据,证明了明确的岩石类型/构造产状/地幔端元的歧视。因此,这些方法的结合,特别是白化数据的KCA,有助于捕获和讨论各种地球化学系统的数据结构,并提供了Excel程序。
Identifying the data structure including trends and groups/clusters in geochemical problems is essential to discuss the origin of sources and processes from the observed variability of data. An increasing number and high dimensionality of recent geochemical data require efficient and accurate multivariate statistical analysis methods. In this paper, we show the relationship and complementary roles of k‐means cluster analysis (KCA), principal component analysis (PCA), and independent component analysis (ICA) to capture the true data structure. When the data are preprocessed by primary standardization (i.e., with the zero mean and normalized by the standard deviation), KCA and PCA provide essentially the same results, although the former returns the solution in a discretized space. When the data are preprocessed by whitening (i.e., normalized by eigenvalues along the principal components), KCA and ICA may identify a set of independent trends and groups, irrespective of the amplitude (power) of variance. As an example, basalt isotopic compositions have been analyzed with KCA on the whitened data, demonstrating clear rock type/tectonic occurrence/mantle end‐member discrimination. Therefore, the combination of these methods, particularly KCA on whitened data, is useful to capture and discuss the data structure of various geochemical systems, for which an Excel program is provided.