Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection

Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection
复制标题

DOI:
10.18653/v1/2020.acl-main.647
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
Shauli Ravfogel;Yanai Elazar;Hila Gonen;Michael Twiton;Yoav Goldberg
Shauli Ravfogel;Yanai Elazar;Hila Gonen;Michael Twiton;Yoav Goldberg
中科院分区:
其他
文献类型:
--
作者:
Shauli Ravfogel;Yanai Elazar;Hila Gonen;Michael Twiton;Yoav Goldberg

文献摘要

被引文献

相似文献

控制神经表征中编码的各种信息的能力有各种各样的用例,特别是考虑到解释这些模型的挑战。我们提出了迭代零空间投影(INLP),一种新的方法,用于从神经表征中删除信息。我们的方法是基于线性分类器的重复训练,这些分类器预测我们要删除的某个属性,然后将表示投影到零空间上。通过这样做,分类器变得不注意该目标属性,使得很难根据它线性地分离数据。虽然适用于多种用途,但我们对偏见和公平用例进行了评估,并表明我们的方法能够减轻单词嵌入中的偏见,以及在多类分类的设置中增加公平性。
The ability to control for the kinds of information encoded in neural representation has a variety of use cases, especially in light of the challenge of interpreting these models. We present Iterative Null-space Projection (INLP), a novel method for removing information from neural representations. Our method is based on repeated training of linear classifiers that predict a certain property we aim to remove, followed by projection of the representations on their null-space. By doing so, the classifiers become oblivious to that target property, making it hard to linearly separate the data according to it. While applicable for multiple uses, we evaluate our method on bias and fairness use-cases, and show that our method is able to mitigate bias in word embeddings, as well as to increase fairness in a setting of multi-class classification.