Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillation

Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillation
复制标题

DOI:
10.1145/3278721.3278725
复制
发表时间:
2017-10
期刊:
Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society
影响因子:
--
通讯作者:
S. Tan;R. Caruana;G. Hooker;Yin Lou
S. Tan;R. Caruana;G. Hooker;Yin Lou
中科院分区:
其他
文献类型:
--
作者:
S. Tan;R. Caruana;G. Hooker;Yin Lou

文献摘要

被引文献

相似文献

黑箱风险评分模型充斥着我们的生活,但通常是专有或不透明的。我们提出了“提炼与比较”(Distill-and-Compare)方法,这是一种在不探究黑箱模型应用程序接口(API)或预先定义要审核的特征的情况下审核此类模型的方法。为了深入了解黑箱模型,我们将它们视为教师,训练透明的学生模型来模拟黑箱模型给出的风险评分。我们将通过提炼训练的模拟模型与在真实结果上训练的第二个未提炼的透明模型进行比较,并利用这两个模型之间的差异来深入了解黑箱模型。我们在四个数据集上展示了该方法:COMPAS、拦截搜身(Stop-and-Frisk)、芝加哥警方以及借贷俱乐部(Lending Club)。我们还提出了一种统计检验方法,以确定一个数据集是否缺失用于训练黑箱模型的关键特征。我们的检验发现,ProPublica的数据可能缺失COMPAS中使用的关键特征。
Black-box risk scoring models permeate our lives, yet are typically proprietary or opaque. We propose Distill-and-Compare, an approach to audit such models without probing the black-box model API or pre-defining features to audit. To gain insight into black-box models, we treat them as teachers, training transparent student models to mimic the risk scores assigned by the black-box models. We compare the mimic model trained with distillation to a second, un-distilled transparent model trained on ground truth outcomes, and use differences between the two models to gain insight into the black-box model. We demonstrate the approach on four data sets: COMPAS, Stop-and-Frisk, Chicago Police, and Lending Club. We also propose a statistical test to determine if a data set is missing key features used to train the black-box model. Our test finds that the ProPublica data is likely missing key feature(s) used in COMPAS.