Making AI Forget You: Data Deletion in Machine Learning

Making AI Forget You: Data Deletion in Machine Learning
复制标题

DOI:
--
复制
发表时间:
2019-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Antonio A. Ginart;M. Guan;G. Valiant;James Y. Zou
Antonio A. Ginart;M. Guan;G. Valiant;James Y. Zou
中科院分区:
其他
文献类型:
--
作者:
Antonio A. Ginart;M. Guan;G. Valiant;James Y. Zou

文献摘要

被引文献

相似文献

最近的激烈讨论集中在如何让个人控制其数据何时可以使用和不可以使用——欧盟的“被遗忘权”法规就是这一努力的一个例子。在本文中,我们启动了一个框架,研究当不再允许部署源自特定用户数据的模型时该怎么办。特别是,我们提出了从训练有素的机器学习模型中有效删除单个数据点的问题。对于许多标准机器学习模型来说,完全删除个人数据的唯一方法是在剩余数据上从头开始重新训练整个模型,这在计算上通常不切实际。我们研究了在机器学习中实现高效数据删除的算法原理。对于 k-means 聚类的特定设置,我们提出了两种可证明有效的删除算法,它们在 6 个数据集上的删除效率平均提高了 100 倍以上,同时生成的聚类的统计质量与规范的 k-means++ 基线相当。
Intense recent discussions have focused on how to provide individuals with control over when their data can and cannot be used --- the EU's Right To Be Forgotten regulation is an example of this effort. In this paper we initiate a framework studying what to do when it is no longer permissible to deploy models derivative from specific user data. In particular, we formulate the problem of efficiently deleting individual data points from trained machine learning models. For many standard ML models, the only way to completely remove an individual's data is to retrain the whole model from scratch on the remaining data, which is often not computationally practical. We investigate algorithmic principles that enable efficient data deletion in ML. For the specific setting of k-means clustering, we propose two provably efficient deletion algorithms which achieve an average of over 100X improvement in deletion efficiency across 6 datasets, while producing clusters of comparable statistical quality to a canonical k-means++ baseline.