Fast and More Powerful Selective Inference for Sparse High-order Interaction Model

Fast and More Powerful Selective Inference for Sparse High-order Interaction Model
复制标题

DOI:
10.1609/aaai.v36i9.21238
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Diptesh Das;Vo Nguyen Le Duy;Hiroyuki Hanada;K. Tsuda;I. Takeuchi
Diptesh Das;Vo Nguyen Le Duy;Hiroyuki Hanada;K. Tsuda;I. Takeuchi
中科院分区:
其他
文献类型:
--
作者:
Diptesh Das;Vo Nguyen Le Duy;Hiroyuki Hanada;K. Tsuda;I. Takeuchi

文献摘要

相似文献

自动化的高风险决策,如医疗诊断,需要具有高可解释性和可靠性的模型。我们认为稀疏高阶相互作用模型是一种可解释的、可靠的、具有良好预测能力的模型。然而,由于组合效应的本质高维性,发现统计上显著的高阶相互作用是具有挑战性的。数据驱动建模中的另一个问题是“择优选择”的影响(即选择偏差)。我们的主要贡献是将最近开发的用于选择性推理的参数化规划方法扩展到高阶交互模型。对樱桃树(所有可能的相互作用)进行详尽的搜索可能是令人生畏和不切实际的,即使对于小型问题也是如此。我们引入了一种高效的剪枝策略,并用合成数据和实际数据证明了该方法的计算效率和统计能力。
Automated high-stake decision-making, such as medical diagnosis, requires models with high interpretability and reliability. We consider the sparse high-order interaction model as an interpretable and reliable model with a good prediction ability. However, finding statistically significant high-order interactions is challenging because of the intrinsically high dimensionality of the combinatorial effects. Another problem in data-driven modeling is the effect of ``cherry-picking" (i.e., selection bias). Our main contribution is extending the recently developed parametric programming approach for selective inference to high-order interaction models. An exhaustive search over the cherry tree (all possible interactions) can be daunting and impractical, even for small-sized problems. We introduced an efficient pruning strategy and demonstrated the computational efficiency and statistical power of the proposed method using both synthetic and real data.