Displaying bias in sampling effort of data accessed from biodiversity databases using ignorance maps.

Displaying bias in sampling effort of data accessed from biodiversity databases using ignorance maps.
复制标题

DOI:
10.3897/bdj.3.e5361
复制
发表时间:
2015
影响因子:
1.3
通讯作者:
Ruete A
Ruete A
中科院分区:
环境科学与生态学4区
文献类型:
--
作者:
Ruete A

文献摘要

被引文献

相似文献

主要包括公民科学数据的开放式生物多样性数据库使广泛的用户可以获得时间和空间上广泛的物种观测数据。然而,这些数据也有局限性,其中包括:有利于记录分布的抽样偏差,缺乏调查工作评估,以及缺乏对所有生物分布的覆盖。这些局限性并不总是记录在案,而基于这些数据的任何技术评估或科学研究都应包括对其源数据的不确定性的评价,研究人员在分析中应承认这一信息。这里提出的无知地图是一种重要而简单的方法,可以实现一种工具,不仅可以直观地探索数据的质量,还可以过滤掉不可靠的结果。我提出了简单的算法来显示无知地图作为一种工具,报告的空间分布的偏见和缺乏采样的努力在整个研究区域。无知分数仅基于原始数据表示,以便尽可能依赖最少的假设。因此,不涉及预测或估计。理由是基于这样的假设,即使用物种组作为取样工作的替代品是适当的,因为通过类似方法观察到的整个物种组可能具有类似的偏倚。然后使用简单的算法将原始数据转换为0-1的无知分数,这些分数很容易比较和扩展。由于需要在大数据集上执行计算,因此简单性对于基于Web的生物多样性信息基础设施的实现至关重要。有了这些算法,任何生物多样性信息的基础设施都可以提供通过它们访问的观察结果的高质量报告。用户可以根据研究问题指定参考分类组和时间范围。该工具的潜力在于其算法的简单性和缺乏对偏倚分布的假设,使用户能够自由地根据其特定需求定制分析。
Open-access biodiversity databases including mainly citizen science data make temporally and spatially extensive species’ observation data available to a wide range of users. Such data have limitations however, which include: sampling bias in favour of recorder distribution, lack of survey effort assessment, and lack of coverage of the distribution of all organisms. These limitations are not always recorded, while any technical assessment or scientific research based on such data should include an evaluation of the uncertainty of its source data and researchers should acknowledge this information in their analysis. The here proposed maps of ignorance are a critical and easy way to implement a tool to not only visually explore the quality of the data, but also to filter out unreliable results. I present simple algorithms to display ignorance maps as a tool to report the spatial distribution of the bias and lack of sampling effort across a study region. Ignorance scores are expressed solely based on raw data in order to rely on the fewest assumptions possible. Therefore there is no prediction or estimation involved. The rationale is based on the assumption that it is appropriate to use species groups as a surrogate for sampling effort because it is likely that an entire group of species observed by similar methods will share similar bias. Simple algorithms are then used to transform raw data into ignorance scores scaled 0-1 that are easily comparable and scalable. Because of the need to perform calculations over big datasets, simplicity is crucial for web-based implementations on infrastructures for biodiversity information. With these algorithms, any infrastructure for biodiversity information can offer a quality report of the observations accessed through them. Users can specify a reference taxonomic group and a time frame according to the research question. The potential of this tool lies in the simplicity of its algorithms and in the lack of assumptions made about the bias distribution, giving the user the freedom to tailor analyses to their specific needs.