Machine truth serum: a surprisingly popular approach to improving ensemble methods

Machine truth serum: a surprisingly popular approach to improving ensemble methods
复制标题

DOI:
10.1007/s10994-022-06183-y
复制
发表时间:
2019-09
期刊:
影响因子:
7.5
通讯作者:
Tianyi Luo;Yang Liu
Tianyi Luo;Yang Liu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Tianyi Luo;Yang Liu

文献摘要

相似文献

群众的智慧(Surowiecki)揭示了一个惊人的事实,即来自群众的多数投票答案通常比少数个别专家更准确。在机器学习中也可以观察到同样的情况——集成方法(Dietterich)利用这个想法在各种环境中利用多种机器学习算法,例如监督学习和半监督学习,通过聚合不同算法的预测来获得比单独从任何组成算法获得的预测更好的性能。然而,当所有组成算法的多数答案更有可能是错误的时候,现有的聚合规则就会失效。在本文中,我们将贝叶斯真值血清(Prelec)中提出的“令人惊讶的更受欢迎的答案更有可能是真实答案而不是大多数答案”的思想扩展到通过集成最终预测方法进一步改进的监督分类和通过集成数据增强方法增强的半监督分类(例如MixMatch (Berthelot et al.,))。我们面临的挑战是定义或检测什么时候答案应该被认为是“令人惊讶的”。我们提出了两种机器学习辅助方法,当少数人而不是大多数人在监督和半监督分类问题的设置上都有正确答案时,它们可以揭示真相。我们将提出的方法命名为“机器促真剂”。我们在一组分类任务(图像、文本等)上的实验表明,在集成最终预测步骤(监督)和集成数据增强步骤(半监督)中应用Machine Truth Serum可以进一步提高分类性能。
Wisdom of the crowd(Surowiecki, ) disclosed a striking fact that the majority voting answer from a crowd is usually more accurate than a few individual experts. The same story is observed in machine learning - ensemble methods (Dietterich, ) leverage this idea to exploit multiple machine learning algorithms in various settings e.g., supervised learning and semi-supervised learning to achieve better performance by aggregating the predictions of different algorithms than that obtained from any constituent algorithm alone. Nonetheless, the existing aggregating rule would fail when the majority answer of all the constituent algorithms is more likely to be wrong. In this paper, we extend the idea proposed inBayesian Truth Serum(Prelec, ) that “a surprisingly more popular answer is more likely to be the true answer instead of the majority one” to supervised classification further improved by ensemble final predictions method and semi-supervised classification (e.g., MixMatch (Berthelot et al., )) enhanced by ensemble data augmentations method. The challenge for us is to define or detect when an answer should be considered as being “surprising”. We present two machine learning aided methods which can reveal the truth when the minority instead of majority has the true answer on both settings of supervised and semi-supervised classification problems. We name our proposed method the Machine Truth Serum. Our experiments on a set of classification tasks (image, text, etc.) show that the classification performance can be further improved by applying Machine Truth Serum in the ensemble final predictions step (supervised) and in the ensemble data augmentations step (semi-supervised).