Exploratory Analysis of Provenance Data Using R and the Provenance Package

Exploratory Analysis of Provenance Data Using R and the Provenance Package
复制标题

使用 R 和 Provenance 包对来源数据进行探索性分析

DOI:
10.3390/min9030193
复制
发表时间:
2019
期刊:
影响因子:
2.5
通讯作者:
P. Vermeesch
P. Vermeesch
中科院分区:
地球科学3区
文献类型:
--
作者:
P. Vermeesch

文献摘要

参考文献

被引文献

相似文献

利用各种化学、矿物学和同位素指标可以追踪碎屑硅质沉积物的来源。这些数据定义了三种不同的数据类型:(1)化学浓度等成分数据;(2)重矿物成分等点计数数据;(3)锆石U-Pb年龄谱等分布数据。这三种数据类型都需要单独的统计处理。任何这种处理的核心是能够量化两个样本之间的“不同之处”。对于成分数据,最好使用经纬度距离。可以使用卡方距离来比较点计数数据,卡方距离比经纬度距离更好地处理缺失分量(零值)。最后,可以使用科尔莫戈罗夫-斯米尔诺夫模型和相关统计方法对分布数据进行比较。对于使用单一种源代理的小型数据集,有时可以通过目测三元图表或年龄谱来进行数据解释。然而,这不再适用于更大、更复杂的数据集。本文回顾了一些多变量排序技术,以帮助解释这些研究。多维尺度(MDS)是一种普遍适用的方法,它将多个样本之间的显著异同表现为相似样本靠近而不相似样本远离的点的配置。对于成分数据,经典的主成分分析与主成分分析是等价的。由此产生的MDS配置可以用组成信息作为双曲线图来扩充。对于点计数数据,卡方距离的经典MDS分析等价于对应分析(CA)。这项技术还可以产生双曲线。因此,MDS提供了一个通用平台来可视化和解释所有类型的来源数据。将该方法推广到三向相异表提供了将几个数据集合并在一起的机会,从而促进了对“大数据”的解释。本文提供了一套使用统计编程语言R的教程。它使用玩具示例说明了成分数据分析、主成分分析、MDS和其他概念的理论基础,然后将这些方法应用于具有种源软件包的实际数据集。
The provenance of siliclastic sediment may be traced using a wide variety of chemical, mineralogical and isotopic proxies. These define three distinct data types: (1) compositional data such as chemical concentrations; (2) point-counting data such as heavy mineral compositions; and (3) distributional data such as zircon U-Pb age spectra. Each of these three data types requires separate statistical treatment. Central to any such treatment is the ability to quantify the `dissimilarity’ between two samples. For compositional data, this is best done using a logratio distance. Point-counting data may be compared using the chi-square distance, which deals better with missing components (zero values) than the logratio distance does. Finally, distributional data can be compared using the Kolmogorov–Smirnov and related statistics. For small datasets using a single provenance proxy, data interpretation can sometimes be done by visual inspection of ternary diagrams or age spectra. However, this no longer works for larger and more complex datasets. This paper reviews a number of multivariate ordination techniques to aid the interpretation of such studies. Multidimensional Scaling (MDS) is a generally applicable method that displays the salient dissimilarities and differences between multiple samples as a configuration of points in which similar samples plot close together and dissimilar samples plot far apart. For compositional data, classical MDS analysis of logratio data is shown to be equivalent to Principal Component Analysis (PCA). The resulting MDS configurations can be augmented with compositional information as biplots. For point-counting data, classical MDS analysis of chi-square distances is shown to be equivalent to Correspondence Analysis (CA). This technique also produces biplots. Thus, MDS provides a common platform to visualise and interpret all types of provenance data. Generalising the method to three-way dissimilarity tables provides an opportunity to combine several datasets together and thereby facilitate the interpretation of `Big Data’. This paper presents a set of tutorials using the statistical programming language R. It illustrates the theoretical underpinnings of compositional data analysis, PCA, MDS and other concepts using toy examples, before applying these methods to real datasets with the provenance package.
DOI: 10.1016/j.epsl.2015.12.036
发表时间: 2016-03
影响因子: 5.3
作者:
M. Rittner;P. Vermeesch;A. Carter;A. Bird;T. Stevens;E. Garzanti;S. Andó;G. Vezzoli;R. Dutt;Zhiwei Xu;Huayu Lu
通讯作者: M. Rittner;P. Vermeesch;A. Carter;A. Bird;T. Stevens;E. Garzanti;S. Andó;G. Vezzoli;R. Dutt;Zhiwei Xu;Huayu Lu
DOI: 10.1016/j.earscirev.2017.11.027
发表时间: 2017-12
影响因子: 12.1
作者:
P. Vermeesch
通讯作者: P. Vermeesch