Improving big citizen science data: Moving beyond haphazard sampling

Improving big citizen science data: Moving beyond haphazard sampling
复制标题

DOI:
10.1371/journal.pbio.3000357
复制
发表时间:
2019-06-01
期刊:
影响因子:
9.8
通讯作者:
Major, Richard E.
Major, Richard E.
中科院分区:
生物学1区
文献类型:
--
作者:
Callaghan, Corey T.;Rowley, Jodi J. L.;Major, Richard E.

文献摘要

被引文献

相似文献

公民科学是主流:每年有数百万人为越来越多的公民科学项目提供数据,形成了庞大的数据集,将在未来几年推动研究。许多公民科学项目实行排行榜框架,根据记录或物种的数量对贡献进行排名,鼓励进一步参与。但是,每个数据点都同样有价值吗?公民科学家收集的数据具有明显的空间和时间偏差,导致不幸的差距和冗余,这给下游分析带来了统计和信息问题。到目前为止,数据的随机结构一直被视为公民科学数据的一个不幸但不可改变的方面。然而,我们认为这个问题实际上是可以解决的:我们提供了一个非常简单、易于处理的框架,可以被大规模的公民科学项目所适应,使公民科学家能够优化他们努力的边际价值,增加整体的集体知识。
Citizen science is mainstream: millions of people contribute data to a growing array of citizen science projects annually, forming massive datasets that will drive research for years to come. Many citizen science projects implement a leaderboard framework, ranking the contributions based on number of records or species, encouraging further participation. But is every data point equally valuable? Citizen scientists collect data with distinct spatial and temporal biases, leading to unfortunate gaps and redundancies, which create statistical and informational problems for downstream analyses. Up to this point, the haphazard structure of the data has been seen as an unfortunate but unchangeable aspect of citizen science data. However, we argue here that this issue can actually be addressed: we provide a very simple, tractable framework that could be adapted by broadscale citizen science projects to allow citizen scientists to optimize the marginal value of their efforts, increasing the overall collective knowledge.