Floating Forests: Quantitative Validation of Citizen Science Data Generated From Consensus Classifications

Floating Forests: Quantitative Validation of Citizen Science Data Generated From Consensus Classifications
复制标题

漂浮森林:对共识分类生成的公民科学数据进行定量验证

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
L. Trouille
L. Trouille
中科院分区:
--
文献类型:
--
作者:
I. Rosenthal;Jarrett E. K. Byrnes;K. Cavanaugh;T. Bell;Briana Harder;A. Haupt;A. Rassweiler;Alejandro P'erez;J. Assis;A. Swanson;Amy Boyer;Adam McMaster;L. Trouille

文献摘要

被引文献

相似文献

大规模的研究工作可能会受到限制可用数据量的后勤限制的阻碍。例如,全球生态问题需要一个全球数据集,而传统的采样协议对于一个小型研究团队来说往往效率太低,无法收集足够数量的数据。公民科学通过众包数据收集提供了另一种选择。尽管越来越受欢迎,但社区接受它的速度很慢,主要是因为担心公民科学家收集的数据质量。使用公民科学项目浮动森林(此http URL),我们表明,由公民科学家的共识分类产生的数据是可比的质量专家生成的分类。浮动森林是一个基于网络的项目,公民科学家在其中查看海岸线的卫星照片并追踪海带斑块的边界。自2014年推出以来,超过7,000名公民科学家已经对加州和塔斯马尼亚的海带森林的750,000多张图像进行了分类。图像由15个用户分类。我们通过覆盖所有公民分类生成共识分类,并通过与专家分类进行比较来评估准确性。计算每个阈值(1-15)的马修斯相关系数(MCC),并认为具有最高MCC的阈值是最佳的。我们发现,最佳用户阈值为4.2,Landsat 5和7的MCC为0.400(0.023 SE),Landsat 8的MCC为0.639(0.246 SE)。这些结果表明,来自共识分类的公民科学数据与专家分类的准确性相当。公民科学项目应实施共识分类等方法,并与专家生成的分类进行定量比较,以避免对数据质量的担忧。
Large-scale research endeavors can be hindered by logistical constraints limiting the amount of available data. For example, global ecological questions require a global dataset, and traditional sampling protocols are often too inefficient for a small research team to collect an adequate amount of data. Citizen science offers an alternative by crowdsourcing data collection. Despite growing popularity, the community has been slow to embrace it largely due to concerns about quality of data collected by citizen scientists. Using the citizen science project Floating Forests (this http URL), we show that consensus classifications made by citizen scientists produce data that is of comparable quality to expert generated classifications. Floating Forests is a web-based project in which citizen scientists view satellite photographs of coastlines and trace the borders of kelp patches. Since launch in 2014, over 7,000 citizen scientists have classified over 750,000 images of kelp forests largely in California and Tasmania. Images are classified by 15 users. We generated consensus classifications by overlaying all citizen classifications and assessed accuracy by comparing to expert classifications. Matthews correlation coefficient (MCC) was calculated for each threshold (1-15), and the threshold with the highest MCC was considered optimal. We showed that optimal user threshold was 4.2 with an MCC of 0.400 (0.023 SE) for Landsats 5 and 7, and a MCC of 0.639 (0.246 SE) for Landsat 8. These results suggest that citizen science data derived from consensus classifications are of comparable accuracy to expert classifications. Citizen science projects should implement methods such as consensus classification in conjunction with a quantitative comparison to expert generated classifications to avoid concerns about data quality.