State Tagging for Improved Earth and Environmental Data Quality Assurance

State Tagging for Improved Earth and Environmental Data Quality Assurance
复制标题

用于改善地球和环境数据质量保证的状态标记

DOI:
10.3389/fenvs.2020.00046
复制
发表时间:
2020
影响因子:
3.7
通讯作者:
J. Watkins
J. Watkins
中科院分区:
经济学3区
文献类型:
--
作者:
M. Tso;P. Henrys;S. Rennie;J. Watkins

文献摘要

被引文献

相似文献

环境数据使我们能够监控我们所生活的不断变化的环境。它使我们能够研究趋势并帮助我们开发更好的模型来描述我们环境中的过程,而它们反过来又可以提供信息来改进管理实践。为了确保数据对于分析和解释来说是可靠的,它们必须经过质量保证程序。此类程序通常包括采样和实验室测量(如果适用)期间的标准操作程序,以及输入数据库时​​的数据验证。后者通常涉及合规性(即格式)和一致性(即值)检查,最有可能采用单参数范围测试的形式。此类测试不考虑进行每次测量时的系统状态,并且向用户提供的有关测量被标记为超出范围的可能原因的上下文信息很少。我们建议使用数据科学技术来用已识别的系统状态来标记每个测量。这里术语“状态”的定义很宽松,它们是使用 k 均值聚类(一种无监督机器学习方法)来识别的。状态的含义可以由专家解释。一旦确定了状态,就可以计算每个观测变量的状态相关预测区间。这种方法为用户提供了更多的上下文信息,以解决超出范围的标志,并得出考虑系统状态变化的观测变量的预测区间。然后,用户可以根据需要应用进一步的分析和过滤。我们用英国两个完善的长期监测数据集来说明我们的方法:来自英国环境变化网络 (ECN) 的飞蛾和蝴蝶数据以及英国 CEH 坎布里亚湖监测计划。我们的工作有助于不断开发更好的数据科学框架,使研究人员和其他利益相关者能够更轻松地查找和使用他们需要的数据。
Environmental data allows us to monitor the constantly changing environment that we live in. It allows us to study trends and helps us to develop better models to describe processes in our environment and they, in turn, can provide information to improve management practices. To ensure that the data are reliable for analysis and interpretation, they must undergo quality assurance procedures. Such procedures generally include standard operating procedures during sampling and laboratory measurement (if applicable), as well as data validation upon entry to databases. The latter usually involves compliance (i.e., format) and conformity (i.e., value) checks that are most likely to be in the form of single parameter range tests. Such tests take no consideration of the system state at which each measurement is made, and provide the user with little contextual information on the probable cause for a measurement to be flagged out of range. We propose the use of data science techniques to tag each measurement with an identified system state. The term “state” here is defined loosely and they are identified using k-means clustering, an unsupervised machine learning method. The meaning of the states is open to specialist interpretation. Once the states are identified, state-dependent prediction intervals can be calculated for each observational variable. This approach provides the user with more contextual information to resolve out-of-range flags and derive prediction intervals for observational variables that considers the changes in system states. The users can then apply further analysis and filtering as they see fit. We illustrate our approach with two well-established long-term monitoring datasets in the UK: moth and butterfly data from the UK Environmental Change Network (ECN), and the UK CEH Cumbrian Lakes monitoring scheme. Our work contributes to the ongoing development of a better data science framework that allows researchers and other stakeholders to find and use the data they need more readily.