Mineral-resource prediction using advanced data analytics and machine learning of the QUEST-South stream-sediment geochemical data, southwestern British Columbia, Canada

Mineral-resource prediction using advanced data analytics and machine learning of the QUEST-South stream-sediment geochemical data, southwestern British Columbia, Canada
复制标题

使用加拿大不列颠哥伦比亚省西南部 QUEST-South 河流沉积物地球化学数据的高级数据分析和机器学习进行矿产资源预测

DOI:
--
复制
发表时间:
2020
期刊:
Geochemistry: Exploration, Environment, Analysis
影响因子:
--
通讯作者:
D. Arne
D. Arne
中科院分区:
--
文献类型:
--
作者:
E. Grunsky;D. Arne

文献摘要

被引文献

相似文献

在这项研究中,我们采用多元统计和预测分类方法来解释地球化学数据从8545河流沉积物样品收集在南不列颠哥伦比亚省,加拿大。对35种元素的数据进行了实验室偏倚校正,并对低于检测下限的报告值进行了调整。每个采样点都与最近的不列颠哥伦比亚省MINFILE发生在2.5公里内。 根据不列颠哥伦比亚省地质调查局矿物存款模型和地球化学特征之间的相似性,MINFILE矿点被分组为“GroupModels”。这些数据用于创建474个观察结果的训练数据集,包括100个未归因于MINFILE事件的样本。训练集被用来生成预测的矿物存款模型,从后验概率估计其余8071个样本。数据进行了集中的对数比变换,然后使用主成分分析(PCA)或t-分布的随机邻居嵌入使用9维(t-SNE)的随机森林分类之前的特征。与使用PCA度量获得的后验概率相比,从t-SNE度量生成的后验概率提供了略高水平的预测准确度。结果是可比的,使用传统的集水区分析方法和专家驱动的模型。这里提出的方法提供了一个可重复的,一致的和防御的方法来识别潜在的矿化地形和矿物系统。
In this study we apply multivariate statistical and predictive classification methods to interpret geochemical data from 8545 stream-sediment samples collected in southern British Columbia, Canada. Data for 35 elements were corrected for laboratory bias and adjusted for values reported below the lower limit of detection. Each sample site was attributed with the closest British Columbia MINFILE occurrence within 2.5 km. MINFILE occurrences were grouped into ‘GroupModels’ based on similarities between the British Columbia Geological Survey mineral deposit models and geochemical signatures. These data were used to create a training dataset of 474 observations, including 100 samples not attributed with a MINFILE occurrence. The training set was used to generate predictions for the mineral deposit models from which posterior probabilities were estimated for the remaining 8071 samples. The data underwent a centred log-ratio transformation and then characterization using either principal component analysis (PCA) or t-distributed stochastic neighbour embedding using 9 dimensions (t-SNE) prior to classification by random forests. The posterior probabilities generated from the t-SNE metric provide a slightly higher level of prediction accuracy compared to the posterior probabilities obtained using the PCA metric. The results are comparable to those obtained using a conventional catchment analysis approach and expert-driven model. The approach presented here provides a repeatable, consistent and defensible methodology for the identification of prospective mineralized terrains and mineral systems.