Interpretable Predictions of Tree-based Ensembles via Actionable Feature Tweaking

Interpretable Predictions of Tree-based Ensembles via Actionable Feature Tweaking
复制标题

通过可操作的特征调整对基于树的集成进行可解释的预测

DOI:
--
复制
发表时间:
2017
期刊:
Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
M. Lalmas
M. Lalmas
中科院分区:
--
文献类型:
--
作者:
Gabriele Tolomei;F. Silvestri;Andrew Haines;M. Lalmas

文献摘要

被引文献

相似文献

机器学习的模型常常被描述为“黑匣子”。然而,在许多现实世界的应用中,模型可能不得不牺牲预测能力,以支持人类的可解释性。在这种情况下,功能工程成为一项关键任务,这需要大量且耗时的人力努力。虽然有些特征本质上是静态的,代表了不受影响的属性(例如,个人的年龄),但另一些特征捕捉到了可以调整的特征(例如,每天摄入的碳水化合物的量)。尽管如此,一旦模型从数据中学习,它对新实例做出的每个预测都是不可逆的-假设每个实例都是位于所选特征空间中的一个静态点。然而,在许多情况下,重要的是要了解(I)为什么模型输出关于给定实例的特定预测,(Ii)应该修改该实例的哪些可调整特征,以及(Iii)当突变实例被输入回模型时如何改变这种预测。在本文中,我们提出了一种技术,该技术利用基于树的集成分类器的内部结构来提供将真实负面实例转换为正面预测实例的建议。我们用一个在线广告应用程序证明了我们方法的有效性。首先,我们设计了一个随机森林分类器,有效地区分了两种类型的广告:低(负面)和高(正面)质量的广告(实例)。然后,我们介绍了一个算法,该算法提供了旨在将低质量广告(负面实例)转换为高质量广告(正面实例)的推荐。最后,我们在一个大型广告网络雅虎双子座的活跃库存的子集上评估了我们的方法。
Machine-learned models are often described as "black boxes". In many real-world applications however, models may have to sacrifice predictive power in favour of human-interpretability. When this is the case, feature engineering becomes a crucial task, which requires significant and time-consuming human effort. Whilst some features are inherently static, representing properties that cannot be influenced (e.g., the age of an individual), others capture characteristics that could be adjusted (e.g., the daily amount of carbohydrates taken). Nonetheless, once a model is learned from the data, each prediction it makes on new instances is irreversible - assuming every instance to be a static point located in the chosen feature space. There are many circumstances however where it is important to understand (i) why a model outputs a certain prediction on a given instance, (ii) which adjustable features of that instance should be modified, and finally (iii) how to alter such a prediction when the mutated instance is input back to the model. In this paper, we present a technique that exploits the internals of a tree-based ensemble classifier to offer recommendations for transforming true negative instances into positively predicted ones. We demonstrate the validity of our approach using an online advertising application. First, we design a Random Forest classifier that effectively separates between two types of ads: low (negative) and high (positive) quality ads (instances). Then, we introduce an algorithm that provides recommendations that aim to transform a low quality ad (negative instance) into a high quality one (positive instance). Finally, we evaluate our approach on a subset of the active inventory of a large ad network, Yahoo Gemini.