Uniform versus uncertainty sampling: When being active is less efficient than staying passive

Uniform versus uncertainty sampling: When being active is less efficient than staying passive
复制标题

均匀采样与不确定性采样:主动的效率低于被动的效率

DOI:
10.48550/arxiv.2212.00772
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Fanny Yang
Fanny Yang
中科院分区:
--
文献类型:
--
作者:
A. Tifrea;Jacob Clarysse;Fanny Yang

文献摘要

参考文献

被引文献

相似文献

人们普遍认为,在相同的标记预算下,像不确定性采样这样的主动学习算法比被动学习(即均匀采样)具有更好的预测性能,尽管计算代价更高。最近的经验证据表明,这种额外的成本可能是徒劳的,因为不确定性抽样有时会比被动学习表现得更差。虽然已有的工作在低维情况下给出了不同的解释,但本文表明,在高维情况下,潜在的机制是完全不同的:对于Logistic回归,我们证明了即使在无噪声数据和使用贝叶斯最优分类器的不确定性的情况下,被动学习的性能也优于不确定性采样。从我们的证明中得到的见解表明,当类之间的间隔较小时,这种高维现象会加剧。我们在20个高维数据集上的实验证实了这一直觉,这些数据集跨越了从金融和组织学到化学和计算机视觉的各种应用范围。
It is widely believed that given the same labeling budget, active learning algorithms like uncertainty sampling achieve better predictive performance than passive learning (i.e. uniform sampling), albeit at a higher computational cost. Recent empirical evidence suggests that this added cost might be in vain, as uncertainty sampling can sometimes perform even worse than passive learning. While existing works offer different explanations in the low-dimensional regime, this paper shows that the underlying mechanism is entirely different in high dimensions: we prove for logistic regression that passive learning outperforms uncertainty sampling even for noiseless data and when using the uncertainty of the Bayes optimal classifier. Insights from our proof indicate that this high-dimensional phenomenon is exacerbated when the separation between the classes is small. We corroborate this intuition with experiments on 20 high-dimensional datasets spanning a diverse range of applications, from finance and histology to chemistry and computer vision.
自我训练将混合模型中的弱学习者转变为强学习者
DOI: --
发表时间: 2022
期刊: International Conference on Artificial Intelligence and Statistics
影响因子: --
作者:
Frei, Spencer;Zou, Difan;Chen, Zixiang;Gu, Quanquan
通讯作者: Gu, Quanquan
DOI: --
发表时间: 2019-03
期刊: ArXiv
影响因子: --
作者:
Mingchen Li;M. Soltanolkotabi;Samet Oymak
通讯作者: Mingchen Li;M. Soltanolkotabi;Samet Oymak
DOI: --
发表时间: 2018-05
期刊: arXiv: Machine Learning
影响因子: --
作者:
Dimitris Tsipras;Shibani Santurkar;Logan Engstrom;Alexander Turner;A. Madry
通讯作者: Dimitris Tsipras;Shibani Santurkar;Logan Engstrom;Alexander Turner;A. Madry
DOI: --
发表时间: 2020-05
期刊: J. Mach. Learn. Res.
影响因子: --
作者:
Vidya Muthukumar;Adhyyan Narang;Vignesh Subramanian;M. Belkin;Daniel J. Hsu;A. Sahai
通讯作者: Vidya Muthukumar;Adhyyan Narang;Vignesh Subramanian;M. Belkin;Daniel J. Hsu;A. Sahai