Using t-distributed Stochastic Neighbor Embedding (t-SNE) for cluster analysis and spatial zone delineation of groundwater geochemistry data

Using t-distributed Stochastic Neighbor Embedding (t-SNE) for cluster analysis and spatial zone delineation of groundwater geochemistry data
复制标题

DOI:
10.1016/j.jhydrol.2021.126146
复制
发表时间:
2021-03
影响因子:
6.4
通讯作者:
Honghua Liu;Jing Yang;M. Ye;S. James;Zhonghua Tang;Jie Dong;Tongju Xing
Honghua Liu;Jing Yang;M. Ye;S. James;Zhonghua Tang;Jie Dong;Tongju Xing
中科院分区:
地球科学1区
文献类型:
--
作者:
Honghua Liu;Jing Yang;M. Ye;S. James;Zhonghua Tang;Jie Dong;Tongju Xing

文献摘要

被引文献

相似文献

聚类分析是了解地下水地球化学的空间和时间模式(例如空间区域)的宝贵工具。为了确定现实问题中未知的聚类数量和聚类成员资格,已经使用了多种方法来辅助聚类分析,其中图形方法是流行且直观的。本研究首次引入t分布随机邻域嵌入(t-SNE)方法作为辅助地下水地球化学数据聚类分析的图形方法。将层次聚类分析(HCA)应用于原始地下水地球化学数据,并使用t-SNE来帮助确定聚类数量和聚类成员资格。随后,t-SNE 被用来帮助圈定地下水地球化学的空间区域。将基于 Thet-SNE 的聚类可视化与基于主成分分析 (PCA) 的可视化进行比较。通过将 HCA、PCA 和 t-SNE 应用于三个地球化学数据集(奥斯陆样线、太原岩溶水和江汉平原地下水数据集,其特点是在不同空间和时间尺度上收集的样本数量和特征不同),我们发现 t-SNE 优于 PCA,可以协助 HCA 作为一种有前景的工具,帮助确定 HCA 簇的数量和划分地下水地球化学的空间区域。应该指出的是,t-SNE 不能单独用于聚类分析,部分原因是 t-SNE 可视化依赖于一个称为困惑度的超参数,该参数对于现实世界的问题是先验未知的。本研究中使用的困惑度值是根据经验确定的,对于具有 14 个样本的太原岩溶水数据集,使用了较小的值 0.1。对于另外两个具有数百个样本的数据集,相应的困惑度值为20和30,在常用的int-SNE的5-50范围内。
Cluster analysis is a valuable tool for understanding spatial and temporal patterns (e.g., spatial zones) of groundwater geochemistry. To determine cluster numbers and cluster memberships that are unknown in real-world problems, a number of methods have been used to assist cluster analysis, among which graphic approaches are popular and intuitive. This study introduced, for the first time, thet-distributed Stochastic Neighbor Embedding (t-SNE) method as a graphic approach to assist cluster analysis for groundwater geochemistry data. The hierarchical cluster analysis (HCA) was applied to original groundwater geochemistry data, andt-SNE was used to help determine the number of cluster and cluster memberships. Afterward,t-SNE was used to help delineate spatial zones of groundwater geochemistry. Thet-SNE-based cluster visualization was compared to the visualization based on principal component analysis (PCA). By applying HCA, PCA, andt-SNE to three geochemical datasets (Oslo transect, Taiyuan karst water, and Jianghan Plain groundwater datasets, which are characterized by different number of samples and features collected across different space and time scales), we found thatt-SNE outperformed PCA to assist HCA as a promising tool for helping determine the number of HCA clusters and delineate spatial zones of groundwater geochemistry. It should be noted thatt-SNE alone cannot be used for cluster analyses, partly becauset-SNE visualization depends on a hyperparameter called perplexity that isa prioriunknown for real-world problems. The perplexity values used in this study were determined empirically, and a small value of 0.1 was used for the Taiyuan karst water dataset with 14 samples. For the other two datasets with hundreds of samples, the corresponding perplexity values were 20 and 30, within the range of 5 – 50 commonly used int-SNE.