Promises and Pitfalls of Black-Box Concept Learning Models

Promises and Pitfalls of Black-Box Concept Learning Models
复制标题

黑盒概念学习模型的前景和陷阱

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
Weiwei Pan
Weiwei Pan
中科院分区:
--
文献类型:
--
作者:
Anita Mahinpei;Justin Clark;Isaac Lage;F. Doshi;Weiwei Pan

文献摘要

被引文献

相似文献

将概念学习作为决策过程中间步骤的机器学习模型可以与黑盒预测模型的性能相匹配,同时保留以人类可理解的术语解释结果的能力。然而,我们证明这些模型学习的概念表示编码的信息超出了预定义的概念,并且自然缓解策略不能完全发挥作用,从而使下游预测的解释产生误导。我们描述了信息泄漏的机制,并建议采取措施减轻其影响。
Machine learning models that incorporate concept learning as an intermediate step in their decision making process can match the performance of black-box predictive models while retaining the ability to explain outcomes in human understandable terms. However, we demonstrate that the concept representations learned by these models encode information beyond the pre-defined concepts, and that natural mitigation strategies do not fully work, rendering the interpretation of the downstream prediction misleading. We describe the mechanism underlying the information leakage and suggest recourse for mitigating its effects.