The Rashomon Importance Distribution: Getting RID of Unstable, Single Model-based Variable Importance

The Rashomon Importance Distribution: Getting RID of Unstable, Single Model-based Variable Importance
复制标题

DOI:
10.48550/arxiv.2309.13775
复制
发表时间:
2023-09
期刊:
ArXiv
影响因子:
--
通讯作者:
J. Donnelly;Srikar Katta;C. Rudin;E. Browne
J. Donnelly;Srikar Katta;C. Rudin;E. Browne
中科院分区:
其他
文献类型:
--
作者:
J. Donnelly;Srikar Katta;C. Rudin;E. Browne

文献摘要

被引文献

相似文献

量化变量的重要性对于回答遗传学、公共政策和医学等领域的高风险问题至关重要。当前的方法通常计算在给定数据集上训练的给定模型的变量重要性。然而,对于给定的数据集,可能有许多模型可以同样很好地解释目标结果;如果不考虑所有可能的解释,不同的研究人员可能会根据相同的数据得出许多相互矛盾但同样有效的结论。此外,即使考虑到给定数据集的所有可能解释,这些见解也可能无法概括,因为并非所有好的解释在合理的数据扰动下都是稳定的。我们提出了一个新的变量重要性框架,该框架量化了所有良好模型集中变量的重要性,并且在整个数据分布中保持稳定。我们的框架非常灵活,可以与大多数现有模型类和全局变量重要性指标集成。我们通过实验证明,我们的框架可以恢复其他方法失败的复杂模拟设置的可变重要性排名。此外,我们表明我们的框架准确地估计了变量对于基础数据分布的真正重要性。我们为我们的估计器提供了一致性和有限样本错误率的理论保证。最后,我们通过一个真实世界的案例研究证明了它的实用性,该研究探索了哪些基因对于预测艾滋病毒感染者的艾滋病毒载量很重要,强调了以前从未研究过的与艾滋病毒相关的重要基因。代码可以在这里找到。
Quantifying variable importance is essential for answering high-stakes questions in fields like genetics, public policy, and medicine. Current methods generally calculate variable importance for a given model trained on a given dataset. However, for a given dataset, there may be many models that explain the target outcome equally well; without accounting for all possible explanations, different researchers may arrive at many conflicting yet equally valid conclusions given the same data. Additionally, even when accounting for all possible explanations for a given dataset, these insights may not generalize because not all good explanations are stable across reasonable data perturbations. We propose a new variable importance framework that quantifies the importance of a variable across the set of all good models and is stable across the data distribution. Our framework is extremely flexible and can be integrated with most existing model classes and global variable importance metrics. We demonstrate through experiments that our framework recovers variable importance rankings for complex simulation setups where other methods fail. Further, we show that our framework accurately estimates the true importance of a variable for the underlying data distribution. We provide theoretical guarantees on the consistency and finite sample error rates for our estimator. Finally, we demonstrate its utility with a real-world case study exploring which genes are important for predicting HIV load in persons with HIV, highlighting an important gene that has not previously been studied in connection with HIV. Code is available here.