Improving language-supervised object detection with linguistic structure analysis

Improving language-supervised object detection with linguistic structure analysis
复制标题

DOI:
10.1109/cvprw59228.2023.00588
复制
发表时间:
2023-06
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
Arushi Rai;Adriana Kovashka
Arushi Rai;Adriana Kovashka
中科院分区:
其他
文献类型:
--
作者:
Arushi Rai;Adriana Kovashka

文献摘要

相似文献

图像监督对象检测通常使用来自人类注释数据集的描述性标题。然而,在野外字幕采取更广泛的语言风格。我们分析了一种普遍存在的语言形式:叙事。我们研究了叙事性和描述性字幕在语言结构和视觉-文本对齐方面的差异,发现我们可以使用词性、修辞结构理论和多模态话语等语言特征对描述性和叙事性字幕进行分类。然后,我们使用它来选择字幕,从中提取图像级标签作为弱监督对象检测的监督。我们还通过基于描述性和叙述性标题与动词类型的接近度进行过滤来提高提取标签的质量。
Language-supervised object detection typically uses descriptive captions from human-annotated datasets. However, in-the-wild captions take on wider styles of language. We analyze one particular ubiquitous form of language: narrative. We study the differences in linguistic structure and visual-text alignment in narrative and descriptive captions and find we can classify descriptive and narrative style captions using linguistic features such as part of speech, rhetoric structure theory, and multimodal discourse. Then, we use this to select captions from which to extract image-level labels as supervision for weakly supervised object detection. We also improve the quality of extracted labels by filtering based on proximity to verb types for both descriptive and narrative captions.