A data-driven method for estimating the composition of end-members from stream water chemistry time series

A data-driven method for estimating the composition of end-members from stream water chemistry time series
复制标题

DOI:
10.5194/hess-26-1977-2022
复制
发表时间:
2022-04
影响因子:
6.3
通讯作者:
Esther Xu Fei;C. Harman
Esther Xu Fei;C. Harman
中科院分区:
地球科学2区
文献类型:
--
作者:
Esther Xu Fei;C. Harman

文献摘要

相似文献

抽象。端元混合分析(EMMA)是一种解释河流水化学变化的方法,广泛用于化学过程线的分离。它是基于这样的假设,即流水是一个保守的混合物,不同的贡献,从良好的特征源解决方案(端元)。这些端元通常通过从流域内收集潜在端元源沃茨的样本并将其与观测结果进行比较来识别。在这里,我们介绍了一个互补的数据驱动的方法(凸船体端元混合分析- CHEMMA)来推断端元的组成和相关的不确定性,从单独的溪流水观测。该方法包括两个步骤。第一个使用凸船体非负矩阵分解(CH-NMF)推断可能的端员组成通过搜索一个简单的,最佳封闭的流水观测。第二步使用约束K均值聚类(COP-KMEANS)对CH-NMF重复应用的结果进行分类,并分析与算法相关的不确定性。在一个应用实例中,利用1986年至1988年帕诺拉山研究流域数据集,CHEMMA是能够鲁棒地再现三个现场测量的端元在以前的研究中发现,只使用流水化学观测。CHEMMA还表明,第四个和第五个端员可以(不太稳健)确定。我们研究的不确定性所产生的非唯一性,这是相关的数据结构,CH-NMF解决方案,并从使用真实的和合成数据的样本数的端员识别。结果表明,当数据集包括包含一个端元的极小贡献的样本时,可以鲁棒地识别混合空间,即,含有来自一个端元的极大贡献的样品不是必需的,但确实降低了端元组成的不确定性。
Abstract. End-member mixing analysis (EMMA) is a method of interpreting stream water chemistry variations and is widely used for chemical hydrograph separation. It is based on the assumption that stream water is a conservative mixture of varying contributions from well-characterized source solutions (end-members). These end-members are typically identified by collecting samples of potential end-member source waters from within the watershed and comparing these to the observations. Here we introduce a complementary data-driven method (convex hull end-member mixing analysis – CHEMMA) to infer the end-member compositions and their associated uncertainties from the stream water observations alone. The method involves two steps. The first uses convex hull nonnegative matrix factorization (CH-NMF) to infer possible end-member compositions by searching for a simplex that optimally encloses the stream water observations. The second step uses constrained K-means clustering (COP-KMEANS) to classify the results from repeated applications of CH-NMF and analyzes the uncertainty associated with the algorithm. In an example application utilizing the 1986 to 1988 Panola Mountain Research Watershed dataset, CHEMMA is able to robustly reproduce the three field-measured end-members found in previous research using only the stream water chemical observations. CHEMMA also suggests that a fourth and a fifth end-member can be (less robustly) identified. We examine uncertainties in end-member identification arising from non-uniqueness, which is related to the data structure, of the CH-NMF solutions, and from the number of samples using both real and synthetic data. The results suggest that the mixing space can be identified robustly when the dataset includes samples that contain extremely small contributions of one end-member, i.e., samples containing extremely large contributions from one end-member are not necessary but do reduce uncertainty about the end-member composition.