Taking a 'Big Data' approach to data quality in a citizen science project.

Taking a 'Big Data' approach to data quality in a citizen science project.
复制标题

DOI:
10.1007/s13280-015-0710-4
复制
发表时间:
2015-11
期刊:
影响因子:
6.5
通讯作者:
Hochachka WM
Hochachka WM
中科院分区:
环境科学与生态学2区
文献类型:
--
作者:
Kelling S;Fink D;La Sorte FA;Johnston A;Bruns NE;Hochachka WM

文献摘要

被引文献

相似文献

来自精心设计的实验的数据为生物多样性研究中的因果关系提供了最有力的证据。然而,对于许多物种来说,这些数据的收集无法扩展到理解种群水平模式所需的空间和时间范围。只有从公民科学项目收集的数据才能收集足够数量的数据,但从志愿者收集的数据本质上是嘈杂和异构的。在这里,我们描述了一种提高 eBird 数据质量的“大数据”方法,eBird 是一个收集鸟类观测数据的全球公民科学项目。首先,eBird的数据提交设计确保所有数据满足完整性和准确性的高标准。其次,我们采用“传感器校准”方法来测量 eBird 参与者检测和识别鸟类能力的个体差异。第三,我们使用物种分布模型来填补数据空白。最后,我们提供了探索鸟类分布中种群水平模式的新颖分析示例。
Data from well-designed experiments provide the strongest evidence of causation in biodiversity studies. However, for many species the collection of these data is not scalable to the spatial and temporal extents required to understand patterns at the population level. Only data collected from citizen science projects can gather sufficient quantities of data, but data collected from volunteers are inherently noisy and heterogeneous. Here we describe a ‘Big Data’ approach to improve the data quality in eBird, a global citizen science project that gathers bird observations. First, eBird’s data submission design ensures that all data meet high standards of completeness and accuracy. Second, we take a ‘sensor calibration’ approach to measure individual variation in eBird participant’s ability to detect and identify birds. Third, we use species distribution models to fill in data gaps. Finally, we provide examples of novel analyses exploring population-level patterns in bird distributions.