Interpretable Companions for Black-Box Models

Interpretable Companions for Black-Box Models
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Dan-qing Pan;Tong Wang;Satoshi Hara
Dan-qing Pan;Tong Wang;Satoshi Hara
中科院分区:
其他
文献类型:
--
作者:
Dan-qing Pan;Tong Wang;Satoshi Hara

文献摘要

相似文献

我们为任何预先训练的黑盒分类器提供了一个可解释的伴随模型。这个想法是,对于任何输入,用户可以决定要么从黑盒模型接收高精度但没有解释的预测,要么采用伴随规则来获得精度稍低的可解释预测。伴随模型是根据数据和黑盒模型的预测进行训练的,目标是结合透明度下的区域——准确率曲线和模型复杂性。我们的模型为面临在总是使用可解释模型和总是使用黑盒模型进行预测任务之间做出选择的困境的从业者提供了灵活的选择,因此,对于任何给定的输入,如果用户发现预测性能令人满意,则可以后退一步,诉诸可解释预测,或者如果规则不满意,则坚持使用黑盒模型。为了展示同伴模型的价值,我们设计了对一百多人的人类评估,以调查可容忍的准确性损失,以获得人类的可解释性。
We present an interpretable companion model for any pre-trained black-box classifiers. The idea is that for any input, a user can decide to either receive a prediction from the black-box model, with high accuracy but no explanations, or employ a companion rule to obtain an interpretable prediction with slightly lower accuracy. The companion model is trained from data and the predictions of the black-box model, with the objective combining area under the transparency--accuracy curve and model complexity. Our model provides flexible choices for practitioners who face the dilemma of choosing between always using interpretable models and always using black-box models for a predictive task, so users can, for any given input, take a step back to resort to an interpretable prediction if they find the predictive performance satisfying, or stick to the black-box model if the rules are unsatisfying. To show the value of companion models, we design a human evaluation on more than a hundred people to investigate the tolerable accuracy loss to gain interpretability for humans.