Deep learning-based diatom taxonomy on virtual slides

Deep learning-based diatom taxonomy on virtual slides
复制标题

DOI:
10.1038/s41598-020-71165-w
复制
发表时间:
2020-09-02
期刊:
影响因子:
4.6
通讯作者:
Nattkemper, Tim W.
Nattkemper, Tim W.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Kloster, Michael;Langenkaemper, Daniel;Nattkemper, Tim W.

文献摘要

被引文献

相似文献

深度卷积神经网络作为图像监督分类的最新方法,也在分类识别的背景下出现。不同的形态和成像技术应用于有机体群体导致高度特定的图像域,这需要定制深度学习解决方案。在这里,我们提供了一个使用深度卷积神经网络(cnn)对形态多样的硅藻微藻群进行分类鉴定的例子。通过结合高分辨率载玻片扫描显微镜、基于网络的协作图像注释和硅藻定制图像分析,我们组装了来自两次南大洋探险的硅藻图像数据库。我们使用这些数据来研究CNN架构、背景掩蔽、数据集大小和可能的概念漂移对图像分类性能的影响。令人惊讶的是,相对较老的网络架构VGG16在我们的图像上表现出了最好的性能和泛化能力。与之前的研究不同,我们发现背景屏蔽略微提高了性能。一般来说,仅在广泛而非特定于领域的图像数据上预训练的卷积层上训练分类器显示出惊人的高性能(F1分数约为97%),每个类的样本数量相对较少(100-300),这表明领域适应新的分类组是可行的,只需投入有限的努力。
Deep convolutional neural networks are emerging as the state of the art method for supervised classification of images also in the context of taxonomic identification. Different morphologies and imaging technologies applied across organismal groups lead to highly specific image domains, which need customization of deep learning solutions. Here we provide an example using deep convolutional neural networks (CNNs) for taxonomic identification of the morphologically diverse microalgal group of diatoms. Using a combination of high-resolution slide scanning microscopy, web-based collaborative image annotation and diatom-tailored image analysis, we assembled a diatom image database from two Southern Ocean expeditions. We use these data to investigate the effect of CNN architecture, background masking, data set size and possible concept drift upon image classification performance. Surprisingly, VGG16, a relatively old network architecture, showed the best performance and generalizing ability on our images. Different from a previous study, we found that background masking slightly improved performance. In general, training only a classifier on top of convolutional layers pre-trained on extensive, but not domain-specific image data showed surprisingly high performance (F1 scores around 97%) with already relatively few (100-300) examples per class, indicating that domain adaptation to a novel taxonomic group can be feasible with a limited investment of effort.