A Latent Variable Model for Discovering Bird Species Commonly Misidentified by Citizen Scientists

A Latent Variable Model for Discovering Bird Species Commonly Misidentified by Citizen Scientists
复制标题

用于发现公民科学家经常错误识别的鸟类的潜在变量模型

DOI:
--
复制
发表时间:
2014
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Weng
Weng
中科院分区:
--
文献类型:
--
作者:
Jun Yu;R. Hutchinson;Weng

文献摘要

被引文献

相似文献

对于像eBird这样的大型公民科学项目来说,数据质量是一个常见的问题。以eBird为例,数据质量差的一个主要原因是缺乏经验的贡献者对鸟类物种的错误识别。提高数据质量的一种积极的方法是发现通常被错误识别的鸟类物种,并教导没有经验的观鸟者这些物种之间的差异。为了实现这一目标,我们开发了一个潜在变量图形模型,该模型可以识别经常被eBird参与者混淆的鸟类种群。我们的模型是生态学文献中经典的占用探测模型的多物种扩展。这种多物种扩展需要一个结构学习步骤和一个计算昂贵的参数学习阶段,我们通过变分近似使其高效。我们发现,我们的模型不仅可以发现错误识别的物种群,而且通过将这些错误识别纳入模型,它还可以实现更准确的物种占用和检测预测。
Data quality is a common source of concern for large-scale citizen science projects like eBird. In the case of eBird, a major cause of poor quality data is the misidentification of bird species by inexperienced contributors. A proactive approach for improving data quality is to discover commonly misidentified bird species and to teach inexperienced birders the differences between these species. To accomplish this goal, we develop a latent variable graphical model that can identify groups of bird species that are often confused for each other by eBird participants. Our model is a multi-species extension of the classic occupancy-detection model in the ecology literature. This multi-species extension requires a structure learning step as well as a computationally expensive parameter learning stage which we make efficient through a variational approximation. We show that our model can not only discover groups of misidentified species, but by including these misidentifications in the model, it can also achieve more accurate predictions of both species occupancy and detection.