Discover the Unknown Biased Attribute of an Image Classifier

Discover the Unknown Biased Attribute of an Image Classifier
复制标题

DOI:
10.1109/iccv48922.2021.01470
复制
发表时间:
2021-04
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Zhiheng Li;Chenliang Xu
Zhiheng Li;Chenliang Xu
中科院分区:
其他
文献类型:
--
作者:
Zhiheng Li;Chenliang Xu

文献摘要

被引文献

相似文献

最近的研究发现,人工智能算法从数据中学习偏差。因此,识别人工智能算法中的偏差是非常迫切和至关重要的。然而,以前的偏见识别管道过度依赖人类专家来猜测潜在的偏见(例如性别),这可能会忽视人类没有意识到的其他潜在的偏见。为了帮助人类专家更好地发现人工智能算法的偏差,本文研究了一个新的问题--对于预测输入图像的目标属性的分类器,发现其未知的有偏属性;为了解决这个具有挑战性的问题,我们使用生成模型的潜在空间中的一个超平面来表示图像属性,从而将原始问题转化为对超平面的法向量和偏移量的优化。在此框架下,我们提出了一种新的全变差损失作为目标函数,并提出了一种新的正交化惩罚作为约束。后者防止发现的偏向属性与目标或已知偏向属性之一相同的平凡解决方案。在解缠数据集和真实数据集上的大量实验表明,该方法可以发现有偏差的属性,并实现更好的解缠w.r.t.目标属性。此外,定性结果表明,该方法可以发现不同对象和场景分类器的不可察觉的偏向属性,从而证明了该方法在不同图像领域中检测偏向属性的普适性。
Recent works find that AI algorithms learn biases from data. Therefore, it is urgent and vital to identify biases in AI algorithms. However, the previous bias identification pipeline overly relies on human experts to conjecture potential biases (e.g., gender), which may neglect other underlying biases not realized by humans. To help human experts better find the AI algorithms’ biases, we study a new problem in this work – for a classifier that predicts a target attribute of the input image, discover its unknown biased attribute.To solve this challenging problem, we use a hyperplane in the generative model’s latent space to represent an image attribute; thus, the original problem is transformed to optimizing the hyperplane’s normal vector and offset. We propose a novel total-variation loss within this framework as the objective function and a new orthogonalization penalty as a constraint. The latter prevents trivial solutions in which the discovered biased attribute is identical with the target or one of the known-biased attributes. Extensive experiments on both disentanglement datasets and real-world datasets show that our method can discover biased attributes and achieve better disentanglement w.r.t. target attributes. Furthermore, the qualitative results show that our method can discover unnoticeable biased attributes for various object and scene classifiers, proving our method’s generalizability for detecting biased attributes in diverse domains of images.