Automated Identification of Herbarium Specimens at Different Taxonomic Levels

Automated Identification of Herbarium Specimens at Different Taxonomic Levels
复制标题

DOI:
10.1007/978-3-319-76445-0_9
复制
发表时间:
2018-01-01
期刊:
MULTIMEDIA TOOLS AND APPLICATIONS FOR ENVIRONMENTAL & BIODIVERSITY INFORMATICS
影响因子:
--
通讯作者:
Bonnet, Pierre
Bonnet, Pierre
中科院分区:
其他
文献类型:
--
作者:
Carranza-Rojas, Jose;Joly, Alexis;Bonnet, Pierre

文献摘要

被引文献

相似文献

地球上开花植物的估计数量约为40万种。为了通过基于图像的自动化方法对所有已知物种进行分类,目前的植物图像数据集必须变得相当大。为了实现这一点,一些作者已经探索了使用植物标本的可能性。随着植物数据集的增长并开始达到数万个类,不平衡的数据集成为一个难题。由于种内和种间的相似性,这导致模型对某些物种不准确。此外,自动工厂识别本质上是分层的。为了解决数据集不平衡的问题,我们需要通过考虑分类学来分类和计算模型损失的方法,例如,通过将物种分组到更高的分类层次。在这项研究中,我们比较了几种架构的自动植物识别,考虑到植物分类,不仅在物种水平上进行分类,但也在更高的水平,如属和科。
The estimated number of flowering plant species on Earth is around 400,000. In order to classify all known species via automated image-based approaches, current datasets of plant images will have to become considerably larger. To achieve this, some authors have explored the possibility of using herbarium sheet images. As the plant datasets grow and start reaching the tens of thousands of classes, unbalanced datasets become a hard problem. This causes models to be inaccurate for certain species due to intra- and inter-specific similarities. Additionally, automatic plant identification is intrinsically hierarchical. In order to tackle this problem of unbalanced datasets, we need ways to classify and calculate the loss of the model by taking into account the taxonomy, for example, by grouping species at higher taxon levels. In this research we compare several architectures for automatic plant identification, taking into account the plant taxonomy to classify not only at the species level, but also at higher levels, such as genus and family.