Rethinking Logic Minimization for Tabular Machine Learning

Rethinking Logic Minimization for Tabular Machine Learning
复制标题

DOI:
10.1109/tai.2022.3224415
复制
发表时间:
2023-10
期刊:
IEEE Transactions on Artificial Intelligence
影响因子:
--
通讯作者:
Litao Qiao;Weijia Wang;S. Dasgupta;Bill Lin
Litao Qiao;Weijia Wang;S. Dasgupta;Bill Lin
中科院分区:
其他
文献类型:
--
作者:
Litao Qiao;Weijia Wang;S. Dasgupta;Bill Lin

文献摘要

被引文献

相似文献

表格数据集可以看作是逻辑函数,可以使用两级逻辑最小化可以简化,以脱节的正常形式产生最小的逻辑公式,进而可以很容易地将其视为用于二进制分类的可解释的决策规则集。但是,使用逻辑最小化进行表格机器学习有两个问题。首先,表格数据集通常包含具有不同类标签的重叠示例,这些示例必须在应用逻辑最小化之前必须解决,因为逻辑最小化具有一致的逻辑函数。其次,即使没有不一致之处,逻辑最小化通常会产生概括较差的复杂模型,因为它完全适合所有数据点,从而导致有害的过度拟合。如何最好地删除培训实例以消除不一致和过度拟合的情况是高度不平凡的。在本文中,我们提出了一个新型的统计框架,用于去除这些训练样本,以便逻辑最小化可以成为制作机器学习的有效方法。使用拟议的方法,我们能够获得可比的绩效,以提高梯度和整体决策树,这是表格学习竞赛中的获胜假设类别,但采用人为理解的决策形式的解释。据《最好的作者知识》,逻辑最小化和可解释的决策规则方法都无法在表格学习问题中实现最先进的表现。
Tabular datasets can be viewed as logic functions that can be simplified using two-level logic minimization to produce minimal logic formulas in disjunctive normal form, which in turn can be readily viewed as an explainable decision rule set for binary classification. However, there are two problems with using logic minimization for tabular machine learning. First, tabular datasets often contain overlapping examples that have different class labels, which have to be resolved before logic minimization can be applied since logic minimization assumes consistent logic functions. Second, even without inconsistencies, logic minimization alone generally produces complex models with poor generalization because it exactly fits all data points, which leads to detrimental overfitting. How best to remove training instances to eliminate inconsistencies and overfitting is highly nontrivial. In this article, we propose a novel statistical framework for removing these training samples so that logic minimization can become an effective approach to tabular machine learning. Using the proposed approach, we are able to obtain comparable performance as gradient boosted and ensemble decision trees, which have been the winning hypothesis classes in tabular learning competitions, but with human-understandable explanations in the form of decision rules. To the best of authors' knowledge, neither logic minimization nor explainable decision rule methods have been able to achieve the state-of-the-art performance before in tabular learning problems.