Exposing and Correcting the Gender Bias in Image Captioning Datasets and Models

Exposing and Correcting the Gender Bias in Image Captioning Datasets and Models
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Shruti Bhargava;David A. Forsyth
Shruti Bhargava;David A. Forsyth
中科院分区:
其他
文献类型:
--
作者:
Shruti Bhargava;David A. Forsyth

文献摘要

被引文献

相似文献

图像字幕的任务隐含地涉及性别识别。然而,由于数据中的性别偏见,图像字幕模型的性别识别受到影响。此外,由于逐词预测,性别活动偏差会影响字幕预测中的其他词,从而导致众所周知的标签偏差问题。在这项工作中,我们研究了COCO字幕数据集中的性别偏见,并表明它不仅来自上下文的性别统计分布,而且来自人类注释者的有缺陷的注释。我们在训练模型中研究这种偏见造成的问题。我们提出了一种技术,以摆脱偏见的任务分为2个子任务:性别中立的图像字幕和性别分类。通过这种脱钩,可以消除性别背景的影响。我们训练了性别中立的图像字幕模型,即使在对与训练数据具有类似偏差的数据集进行评估时,该模型也能提供与性别模型相当的结果。有趣的是,这个模型对没有人类的图像的预测也明显不同于对性别字幕训练的预测。我们使用图像中人物的可用边界框和基于掩码的注释来训练性别分类器。这使我们能够摆脱背景,专注于人来预测性别。通过将性别替换到性别中立的字幕中,我们得到最终的性别预测。我们的预测与使用性别训练的模型具有相似的性能,同时没有性别偏见。最后,我们的主要结果是,在一个反刻板的数据集上,我们的模型优于一个流行的图像字幕模型,该模型是用性别训练的。
The task of image captioning implicitly involves gender identification. However, due to the gender bias in data, gender identification by an image captioning model suffers. Also, the gender-activity bias, owing to the word-by-word prediction, influences other words in the caption prediction, resulting in the well-known problem of label bias. In this work, we investigate gender bias in the COCO captioning dataset and show that it engenders not only from the statistical distribution of genders with contexts but also from the flawed annotation by the human annotators. We look at the issues created by this bias in the trained models. We propose a technique to get rid of the bias by splitting the task into 2 subtasks: gender-neutral image captioning and gender classification. By this decoupling, the gender-context influence can be eradicated. We train the gender-neutral image captioning model, which gives comparable results to a gendered model even when evaluating against a dataset that possesses a similar bias as the training data. Interestingly, the predictions by this model on images with no humans, are also visibly different from the one trained on gendered captions. We train gender classifiers using the available bounding box and mask-based annotations for the person in the image. This allows us to get rid of the context and focus on the person to predict the gender. By substituting the genders into the gender-neutral captions, we get the final gendered predictions. Our predictions achieve similar performance to a model trained with gender, and at the same time are devoid of gender bias. Finally, our main result is that on an anti-stereotypical dataset, our model outperforms a popular image captioning model which is trained with gender.