Promises and Pitfalls of Black-Box Concept Learning Models
Promises and Pitfalls of Black-Box Concept Learning Models
复制标题
黑盒概念学习模型的前景和陷阱
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Weiwei Pan
中科院分区:
文献类型:
--
作者:
Anita Mahinpei;Justin Clark;Isaac Lage;F. Doshi;Weiwei Pan
Machine learning models that incorporate concept learning as an intermediate step in their decision making process can match the performance of black-box predictive models while retaining the ability to explain outcomes in human understandable terms. However, we demonstrate that the concept representations learned by these models encode information beyond the pre-defined concepts, and that natural mitigation strategies do not fully work, rendering the interpretation of the downstream prediction misleading. We describe the mechanism underlying the information leakage and suggest recourse for mitigating its effects.