Measuring agreement among experts in classifying camera images of similar species

Measuring agreement among experts in classifying camera images of similar species
复制标题

DOI:
10.1002/ece3.4567
复制
发表时间:
2018-11-01
影响因子:
2.6
通讯作者:
Hodges, Karen E.
Hodges, Karen E.
中科院分区:
生物学2区
文献类型:
--
作者:
Gooliaff, T. J.;Hodges, Karen E.

文献摘要

被引文献

相似文献

通过公民科学的方式捕捉和征集野生动物图像已经成为生态学研究的常用工具。这些研究收集了许多野生动物的图像,正确的物种分类是至关重要的;即使是很低的错误分类率也会导致对物种的地理范围或栖息地使用的错误估计,从而可能阻碍保护或管理工作。然而,有些物种很难区分,这使得物种分类具有挑战性,但是专家之间关于分类一致性的文献仍然很少。在这里,我们衡量专家在区分两个相似的同属物种,山猫(山猫鲁弗斯)和加拿大猞猁(加拿大猞猁)的图像上的一致性。我们请专家对选定图像中的物种进行分类,以测试季节、背景栖息地、一天中的时间和每只动物的可见特征(如脸、腿、尾巴)是否会影响专家对每张图像中物种的一致看法。总体而言,专家有中等程度的同意(Fleiss kappa = 0.64),但专家有不同程度的同意取决于这些图像特征。大多数图像(71%)有>= 1专家分类为“未知”,许多图像(39%)有一些专家将图像分类为“山猫”,而另一些则将其分类为“猞猁”。此外,专家们甚至自己也不一致,当他们被要求在几个月后重新分类相同的图像时,他们改变了对许多图像的分类。这些结果表明,单个专家对相似物种的图像分类是不可靠的。大多数图像确实从专家那里获得了明确的多数分类,尽管我们强调,即使大多数分类也可能是错误的。我们建议使用野生动物图像的研究人员咨询多个物种专家,以增加他们对相似同域物种的图像分类的信心。尽管如此,当一个具有相似同域分布的物种的存在必须是决定性的,物理或遗传证据应该是必需的。
Camera trapping and solicitation of wildlife images through citizen science have become common tools in ecological research. Such studies collect many wildlife images for which correct species classification is crucial; even low misclassification rates can result in erroneous estimation of the geographic range or habitat use of a species, potentially hindering conservation or management efforts. However, some species are difficult to tell apart, making species classification challenging-but the literature on classification agreement rates among experts remains sparse. Here, we measure agreement among experts in distinguishing between images of two similar congeneric species, bobcats (Lynx rufus) and Canada lynx (Lynx canadensis). We asked experts to classify the species in selected images to test whether the season, background habitat, time of day, and the visible features of each animal (e.g., face, legs, tail) affected agreement among experts about the species in each image. Overall, experts had moderate agreement (Fleiss' kappa = 0.64), but experts had varying levels of agreement depending on these image characteristics. Most images (71%) had >= 1 expert classification of "unknown," and many images (39%) had some experts classify the image as "bobcat" while others classified it as "lynx." Further, experts were inconsistent even with themselves, changing their classifications of numerous images when they were asked to reclassify the same images months later. These results suggest that classification of images by a single expert is unreliable for similar-looking species. Most of the images did obtain a clear majority classification from the experts, although we emphasize that even majority classifications may be incorrect. We recommend that researchers using wildlife images consult multiple species experts to increase confidence in their image classifications of similar sympatric species. Still, when the presence of a species with similar sympatrics must be conclusive, physical or genetic evidence should be required.