Implementing GitHub Actions continuous integration to reduce error rates in ecological data collection

Implementing GitHub Actions continuous integration to reduce error rates in ecological data collection
复制标题

DOI:
10.1111/2041-210x.13982
复制
发表时间:
2022-09
影响因子:
6.6
通讯作者:
Albert Y. Kim;Valentine Herrmann;Ross Bareto;Brian Calkins;E. Gonzalez-Akre;Daniel J. Johnson;Jennifer A. Jordan;L. Magee;I. McGregor;Nicolle Montero;Karl Novak;Teagan Rogers;J. Shue;K. Anderson‐Teixeira
Albert Y. Kim;Valentine Herrmann;Ross Bareto;Brian Calkins;E. Gonzalez-Akre;Daniel J. Johnson;Jennifer A. Jordan;L. Magee;I. McGregor;Nicolle Montero;Karl Novak;Teagan Rogers;J. Shue;K. Anderson‐Teixeira
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Albert Y. Kim;Valentine Herrmann;Ross Bareto;Brian Calkins;E. Gonzalez-Akre;Daniel J. Johnson;Jennifer A. Jordan;L. Magee;I. McGregor;Nicolle Montero;Karl Novak;Teagan Rogers;J. Shue;K. Anderson‐Teixeira

文献摘要

相似文献

准确的野外数据对于理解生态系统和预测其对全球变化的反应至关重要。然而,数据收集错误是常见的,数据分析往往远远落后于它的收集,许多错误不能再纠正,也不能重新审视异常的观察。需要的是一个系统,其中数据质量保证和控制(QA/QC),以及基本数据摘要的生产,可以在数据收集后立即自动化。在这里,我们实现并测试一个系统来满足这些需求。对于两个森林研究站点的两次年度树木死亡率普查和树木计带调查,我们使用GitHub Actions持续集成(CI)来自动化数据QA/QC,并运行常规数据整理脚本,以生成准备用于分析的清洁数据集。这种系统自动化有许多好处,包括:(1)产生关于数据收集状态和需要纠正的错误的近实时信息,从而使最终的数据集没有可检测到的错误;(2)在现场技术人员中有明显的学习效果,其中现场数据收集的原始错误率在系统实施后显着下降;(3)保证了计算的可重复性-即,系统对代码、数据和软件变化的健壮性。通过实施CI,研究人员可以确保数据集没有任何可以编码测试的错误。其结果是显著提高了数据质量,提高了现场技术人员的技能,减少了对专家监督的需求。此外,我们认为CI的实施是迈向数据收集和分析管道的第一步,该管道也更能响应快速变化的生态动态,使其更适合于在当前环境快速变化的时代研究生态系统。
Accurate field data are essential to understanding ecological systems and forecasting their responses to global change. Yet, data collection errors are common, and data analysis often lags far enough behind its collection that many errors can no longer be corrected, nor can anomalous observations be revisited. Needed is a system in which data quality assurance and control (QA/QC), along with the production of basic data summaries, can be automated immediately following data collection. Here, we implement and test a system to satisfy these needs. For two annual tree mortality censuses and a dendrometer band survey at two forest research sites, we used GitHub Actions continuous integration (CI) to automate data QA/QC and run routine data wrangling scripts to produce cleaned datasets ready for analysis. This system automation had numerous benefits, including (1) the production of near real‐time information on data collection status and errors requiring correction, resulting in final datasets free of detectable errors, (2) an apparent learning effect among field technicians, wherein original error rates in field data collection declined significantly following implementation of the system, and (3) an assurance of computational reproducibility—that is, robustness of the system to changes in code, data and software. By implementing CI, researchers can ensure that datasets are free of any errors for which a test can be coded. The result is dramatically improved data quality, increased skill among field technicians, and reduced need for expert oversight. Furthermore, we view CI implementation as a first step towards a data collection and analysis pipeline that is also more responsive to rapidly changing ecological dynamics, making it better suited to study ecological systems in the current era of rapid environmental change.