Methods of training set construction: Towards improving performance for automated mesozooplankton image classification systems

Methods of training set construction: Towards improving performance for automated mesozooplankton image classification systems
复制标题

DOI:
10.1016/j.csr.2012.01.005
复制
发表时间:
2012-03-15
影响因子:
2.3
通讯作者:
Hsieh, Chih-hao
Hsieh, Chih-hao
中科院分区:
地球科学3区
文献类型:
--
作者:
Chang, Chun-Yi;Ho, Pei-Chi;Hsieh, Chih-hao

文献摘要

被引文献

相似文献

水柱的物理化学性质的变化与浮游动物群落的分类组成之间的对应关系是海洋系统长期和广泛变化的一个重要指标。评估组成变化并将其与各种形式的扰动联系起来,需要能够快速而准确地应用的常规分类鉴定方法。由人类专家进行传统的识别是准确的,但非常耗时。浮游生物群落自动图像分类系统的应用已经成为解决这一限制的潜在方案。这项研究的目的是评估ZooScan系统训练集构建的特定方面如何影响我们将浮游动物分类组成的变化与东中国海水文特性变化联系起来的能力。具体地说,我们比较了浮游动物分类器与以下训练的相对效用:(I)特定水团训练集和全局训练集:(Ii)平衡训练集与不平衡训练集。水团特定分类器的分类性能(准确度和精确度)随着环境的不同而下降,这表明水团特定的分类性能。然而,通过用代表所有水文亚区的样本(即全局分类器)来训练我们的系统,也可以获得类似的分类性能。在检查了特定类别的准确性后,我们发现,由于准确性主要由优势类群决定,因此出现了相同的表现。这种明显较高的分类精度是以对稀有类群的准确分类为代价的。为了探索这种有偏见的分类的基础,我们用每个类别相同数量的训练数据来训练我们的全局分类器(平衡训练)。我们发现,均衡训练在识别稀有类群时具有较高的准确率,但在识别丰富类群时准确率较低。识别中引入的错误仍然对自动分类系统构成重大挑战。为了完全自动化对浮游动物群落的分析,并将组成变化与水文特性联系起来,该系统的识别能力需要进一步改进。(C)2012爱思唯尔有限公司。保留所有权利。
The correspondence between variation in the physico-chemical properties of the water column and the taxonomic composition of zooplankton communities represents an important indicator of long-term and broad-scale change in marine systems. Evaluating and relating compositional change to various forms of perturbation demand routine taxonomic identification methods that can be applied rapidly and accurately. Traditional identification by human experts is accurate but very time-consuming. The application of automated image classification systems for plankton communities has emerged as a potential resolution to this limitation. The objective of this study is to evaluate how specific aspects of training set construction for the ZooScan system influenced our ability to relate variation in zooplankton taxonomic composition to variation of hydrographic properties in the East China Sea. Specifically, we compared the relative utility of zooplankton classifiers trained with the following: (i) water mass-specific and global training sets: (ii) balanced versus imbalanced training sets. The classification performance (accuracy and precision) of water-mass specific classifiers tended to decline with environmental dissimilarity, suggesting water-mass specificity However, similar classification performance was also achieved by training our system with samples representing all hydrographic subregions (i.e. a global classifier). After examining category-specific accuracy, we found that equal performance arises because the accuracy was mainly determined by dominant taxa. This apparently high classification accuracy was at the expense of accurate classification of rare taxa. To explore the basis for such biased classification, we trained our global classifier with an equal amount of training data for each category (balanced training). We found that balanced training had higher accuracy at recognizing rare taxa but low accuracy at abundant taxa. The errors introduced in recognition still pose a major challenge for automatic classification systems. In order to fully automate analyses of zooplankton communities and relate variation in composition to hydrographic properties, the recognition power of the system requires further improvements. (C) 2012 Elsevier Ltd. All rights reserved.