Algorithmic Collective Action in Machine Learning

Algorithmic Collective Action in Machine Learning
复制标题

DOI:
10.48550/arxiv.2302.04262
复制
发表时间:
2023-02
期刊:
--
影响因子:
--
通讯作者:
Moritz Hardt;Eric V. Mazumdar;Celestine Mendler-Dunner;Tijana Zrnic
Moritz Hardt;Eric V. Mazumdar;Celestine Mendler-Dunner;Tijana Zrnic
中科院分区:
其他
文献类型:
--
作者:
Moritz Hardt;Eric V. Mazumdar;Celestine Mendler-Dunner;Tijana Zrnic

文献摘要

被引文献

相似文献

我们在部署机器学习算法的数字平台上启动了算法集体行动的原则性研究。我们提出了一个简单的集体与公司学习算法相互作用的理论模型。集体汇集参与个体的数据,并通过指导参与者如何修改自己的数据来实现集体目标来执行算法策略。我们调查的后果,这个模型在三个基本的学习理论设置:一个非参数的最优学习算法,参数风险最小化,和基于梯度的优化的情况下。在每种情况下,我们提出了协调的算法策略,并将自然成功标准描述为集体规模的函数。为了补充我们的理论,我们对一项技能分类任务进行了系统的实验,该任务涉及来自自由职业者工作平台的数万份简历。通过对BERT类语言模型的2000多次模型训练,我们看到我们的经验观察与我们的理论预测之间出现了惊人的一致性。综上所述,我们的理论和实验广泛支持这样的结论,即极小分数大小的算法集体可以对平台的学习算法施加显着的控制。
We initiate a principled study of algorithmic collective action on digital platforms that deploy machine learning algorithms. We propose a simple theoretical model of a collective interacting with a firm's learning algorithm. The collective pools the data of participating individuals and executes an algorithmic strategy by instructing participants how to modify their own data to achieve a collective goal. We investigate the consequences of this model in three fundamental learning-theoretic settings: the case of a nonparametric optimal learning algorithm, a parametric risk minimizer, and gradient-based optimization. In each setting, we come up with coordinated algorithmic strategies and characterize natural success criteria as a function of the collective's size. Complementing our theory, we conduct systematic experiments on a skill classification task involving tens of thousands of resumes from a gig platform for freelancers. Through more than two thousand model training runs of a BERT-like language model, we see a striking correspondence emerge between our empirical observations and the predictions made by our theory. Taken together, our theory and experiments broadly support the conclusion that algorithmic collectives of exceedingly small fractional size can exert significant control over a platform's learning algorithm.