Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification

Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification
复制标题

DOI:
10.1109/cvpr52729.2023.01839
复制
发表时间:
2022-11
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yue Yang;Artemis Panagopoulou;Shenghao Zhou;Daniel Jin;Chris Callison-Burch;Mark Yatskar
Yue Yang;Artemis Panagopoulou;Shenghao Zhou;Daniel Jin;Chris Callison-Burch;Mark Yatskar
中科院分区:
其他
文献类型:
--
作者:
Yue Yang;Artemis Panagopoulou;Shenghao Zhou;Daniel Jin;Chris Callison-Burch;Mark Yatskar

文献摘要

相似文献

概念瓶颈模型(CBM)是固有的可解释模型,它将模型决策转化为人类可读的概念。它们允许人们很容易地理解模型失败的原因,这是高风险应用程序的一个关键特性。CBMs需要手动指定概念,并且通常表现不如黑盒对应的概念,从而阻碍了它们的广泛采用。我们解决了这些缺点,并首先展示了如何在没有与黑盒模型相似精度的手动规范的情况下构建高性能的cbm。我们的方法,语言引导瓶颈(LaBo),利用语言模型GPT-3来定义可能的瓶颈的大空间。给定一个问题领域,LaBo使用GPT-3生成关于类别的事实句子来形成候选概念。LaBo通过一种新颖的子模块实用程序有效地搜索可能的瓶颈,该实用程序促进了判别和多样化信息的选择。最终,GPT-3的句子概念可以使用CLIP与图像对齐,形成瓶颈层。实验证明LaBo对于视觉识别的重要概念是一个非常有效的先验。在对11个不同数据集的评估中,LaBo瓶颈在少量数据集分类方面表现出色:它们比黑箱线性探针在一次数据集和更多数据集上的准确率高出11.7%。总的来说,LaBo证明了固有的可解释模型可以广泛地应用于与黑盒方法相似或更好的性能。代码和数据可在https://github.com/YueYANG1996/LaBo上获得
Concept Bottleneck Models (CBM) are inherently interpretable models that factor model decisions into humanreadable concepts. They allow people to easily understand why a model is failing, a critical feature for high-stakes applications. CBMs require manually specified concepts and often under-perform their black box counterparts, preventing their broad adoption. We address these shortcomings and are first to show how to construct high-performance CBMs without manual specification of similar accuracy to black box models. Our approach, Language Guided Bottlenecks (LaBo), leverages a language model, GPT-3, to define a large space of possible bottlenecks. Given a problem domain, LaBo uses GPT-3 to produce factual sentences about categories to form candidate concepts. LaBo efficiently searches possible bottlenecks through a novel submodular utility that promotes the selection of discriminative and diverse information. Ultimately, GPT-3's sentential concepts can be aligned to images using CLIP, to form a bottleneck layer. Experiments demonstrate that LaBo is a highly effective prior for concepts important to visual recognition. In the evaluation with 11 diverse datasets, LaBo bottlenecks excel at few-shot classification: they are 11.7% more accurate than black box linear probes at 1 shot and comparable with more data. Overall, LaBo demonstrates that inherently interpretable models can be widely applied at similar, or better, performance than black box approaches.11Code and data are available at https://github.com/YueYANG1996/LaBo