Adapting Vision Foundation Models for Plant Phenotyping

Adapting Vision Foundation Models for Plant Phenotyping
复制标题

DOI:
10.1109/iccvw60793.2023.00067
复制
发表时间:
2023-10
期刊:
2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)
影响因子:
--
通讯作者:
Feng Chen;M. Giuffrida;S. Tsaftaris
Feng Chen;M. Giuffrida;S. Tsaftaris
中科院分区:
其他
文献类型:
--
作者:
Feng Chen;M. Giuffrida;S. Tsaftaris

文献摘要

相似文献

基础模型是在大量数据上预先训练的大型模型。它们通常可以以最小的努力适应不同的下游任务。然而,由于基础模型通常是在来自互联网的图像或文本上进行预训练的,因此它们在植物表型分析等专业领域的性能受到质疑。此外,完全微调基础模型是耗时的,需要高计算能力。本文研究了植物表型设置和任务的基础模型的有效适应。我们进行了广泛的实验微调三个基础模型,MAE,DINO和DINOv2三个基本的植物表型任务:叶片计数,实例分割和疾病分类。特别是,预训练的骨干保持冻结,同时评估两种不同的微调方法,即适配器调优(使用LoRA)和解码器调优。实验结果表明,基础模型可以有效地适应多个植物表型任务,产生与为每个任务专门设计或训练的最新技术(SoTA)模型相似的性能。尽管在不同的任务上表现出很大的可移植性,但在某些情况下,经过微调的基础模型的性能略差于SoTA特定任务模型,这需要进一步研究。
Foundation models are large models pre-trained on tremendous amount of data. They can be typically adapted to diverse downstream tasks with minimal effort. However, as foundation models are usually pre-trained on images or texts sourced from the Internet, their performance in specialized domains, such as plant phenotyping, comes into question. In addition, fully fine-tuning foundation models is time-consuming and requires high computational power. This paper investigates the efficient adaptation of foundation models for plant phenotyping settings and tasks. We perform extensive experiments on fine-tuning three foundation models, MAE, DINO, and DINOv2 on three essential plant phenotyping tasks: leaf counting, instance segmentation, and disease classification. In particular, the pretrained backbones are kept frozen, while two distinct fine-tuning methods are evaluated, namely adapter tuning (using LoRA) and decoder tuning. The experimental results show that a foundation model can be efficiently adapted to multiple plant phenotyping tasks, yielding similar performance as the state-of-the-art (SoTA) models specifically designed or trained for each task. Despite exhibiting great transferability over different tasks, the fine-tuned foundation models perform slightly worse than the SoTA task-specific models in some scenarios, which requires further investigation.