Information-theoretic Classification Accuracy: A Criterion that Guides Data-driven Combination of Ambiguous Outcome Labels in Multi-class Classification

Information-theoretic Classification Accuracy: A Criterion that Guides Data-driven Combination of Ambiguous Outcome Labels in Multi-class Classification
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Chihao Zhang;Y. Chen;Shihua Zhang;Jingyi Jessica Li
Chihao Zhang;Y. Chen;Shihua Zhang;Jingyi Jessica Li
中科院分区:
其他
文献类型:
--
作者:
Chihao Zhang;Y. Chen;Shihua Zhang;Jingyi Jessica Li

文献摘要

相似文献

结果标签的模糊性和主观性在现实世界的数据集中是普遍存在的。虽然从业者通常以特别的方式将所有数据点(实例)的模糊结果标签组合联合收割机以提高多类分类的准确性,但是缺乏通过任何最优性标准来指导所有数据点的标签组合的原则性方法。为了解决这个问题,我们提出了信息理论分类精度(ITCA),一个标准,平衡预测精度(预测标签与实际标签的一致性如何)和分类分辨率(有多少标签是可预测的)之间的权衡,以指导从业者如何联合收割机模糊的结果标签。为了找到ITCA表示的最佳标签组合,我们提出了两种搜索策略:贪婪搜索和广度优先搜索。值得注意的是,ITCA和两种搜索策略适用于所有机器学习分类算法。结合分类算法和搜索策略,ITCA有两个用途:提高预测精度和识别模糊标签。我们首先验证了ITCA实现了高精度与两个搜索策略,找到正确的标签组合的合成和真实的数据。然后,我们证明了ITCA在不同应用中的有效性,包括医疗预后,癌症生存预测,用户人口统计预测和细胞类型分类。我们还提供了理论上的见解ITCA通过研究的预言和线性判别分析分类算法。Python包itca(可在https://github.com/JSB-UCLA/ITCA获得)实现了ITCA和搜索策略。
Outcome labeling ambiguity and subjectivity are ubiquitous in real-world datasets. While practitioners commonly combine ambiguous outcome labels for all data points (instances) in an ad hoc way to improve the accuracy of multi-class classification, there lacks a principled approach to guide the label combination for all data points by any optimality criterion. To address this problem, we propose the information-theoretic classification accuracy (ITCA), a criterion that balances the trade-off between prediction accuracy (how well do predicted labels agree with actual labels) and classification resolution (how many labels are predictable), to guide practitioners on how to combine ambiguous outcome labels. To find the optimal label combination indicated by ITCA, we propose two search strategies: greedy search and breadth-first search. Notably, ITCA and the two search strategies are adaptive to all machine-learning classification algorithms. Coupled with a classification algorithm and a search strategy, ITCA has two uses: improving prediction accuracy and identifying ambiguous labels. We first verify that ITCA achieves high accuracy with both search strategies in finding the correct label combinations on synthetic and real data. Then we demonstrate the effectiveness of ITCA in diverse applications including medical prognosis, cancer survival prediction, user demographics prediction, and cell type classification. We also provide theoretical insights into ITCA by studying the oracle and the linear discriminant analysis classification algorithms. Python package itca (available at https://github.com/JSB-UCLA/ITCA) implements ITCA and search strategies.