Overlearning Reveals Sensitive Attributes

Overlearning Reveals Sensitive Attributes
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Congzheng Song;Vitaly Shmatikov
Congzheng Song;Vitaly Shmatikov
中科院分区:
其他
文献类型:
--
作者:
Congzheng Song;Vitaly Shmatikov

文献摘要

相似文献

“过度学习”意味着为看似简单的目标训练的模型隐含地学习识别以下属性和概念:(1)不是学习目标的一部分,以及(2)从隐私或偏见的角度来看是敏感的。例如,面部图像的二进制性别分类器还学习识别种族,甚至是在训练数据和身份中没有表示的种族。我们在几个视觉和NLP模型中展示了过度学习,并分析了其有害后果。首先,过度学习模型的推理时间表示揭示了输入的敏感属性,破坏了模型划分等隐私保护。其次,即使在没有原始训练数据的情况下,过度学习的模型也可以被“重新利用”用于不同的、侵犯隐私的任务。我们表明,过度学习是内在的一些任务,不能通过审查不需要的属性。最后,我们研究了在模型训练过程中过度学习发生的位置、时间和原因。
"Overlearning" means that a model trained for a seemingly simple objective implicitly learns to recognize attributes and concepts that are (1) not part of the learning objective, and (2) sensitive from a privacy or bias perspective. For example, a binary gender classifier of facial images also learns to recognize races\textemdash even races that are not represented in the training data\textemdash and identities. We demonstrate overlearning in several vision and NLP models and analyze its harmful consequences. First, inference-time representations of an overlearned model reveal sensitive attributes of the input, breaking privacy protections such as model partitioning. Second, an overlearned model can be "re-purposed" for a different, privacy-violating task even in the absence of the original training data. We show that overlearning is intrinsic for some tasks and cannot be prevented by censoring unwanted attributes. Finally, we investigate where, when, and why overlearning happens during model training.