Turning an enemy into an ally: Privacy In Machine Learning (Pri-ML)
Turning an enemy into an ally: Privacy In Machine Learning (Pri-ML)
批准号:
RGPIN-2022-03721
负责人:
Park, MiJung
金额:
$1.82万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
数据具有巨大的潜力,可以为全球饥饿和气候变化等紧迫问题提供改变世界的解决方案。然而,以解决大问题所需的规模共享数据也伴随着隐私问题。我的长期目标是提供方法论基础,使我们能够在不侵犯隐私的情况下正确使用和分享这些潜在的有用数据。由于其可证明性,隐私保护机器学习的最新进展一直围绕着差异隐私。然而,保护隐私需要在机器学习系统中注入噪音,导致隐私和准确性之间的根本权衡。当算法的输出对其输入的任何变化“敏感”时,这种权衡会变得更糟,因为必须增加噪声幅度以掩盖这种变化。在研究方案的A部分,我提出了可以有效限制敏感性的新方法。第二个挑战源于差异隐私的可组合性:每次访问数据都会降低隐私保障。差异私有数据生成通过创建可无限制重复使用的合成数据集来解决这一问题。与现有工作假设某些数据分布或特定的下游任务限制生成数据的有用性不同,在B部分中,我提出了一个高度灵活的基于非参数核距离的框架来比较数据和合成数据分布。第三个挑战出现在未来的算法设计要求上,如一般数据保护法规。在C部分,我提出了满足不同隐私和其他新兴概念、可解释性、因果关系和公平性的新框架。我们对它们相互作用的理论和定量理解将有助于开发同时考虑这些概念的实用算法。该研究项目有两个重要的社会影响。首先,像B部分那样的技术确保了隐私保护,这可以促进更多的数据共享,造福于公共利益。其次,正如Netflix纪录片Coded Bias所指出的那样,基于机器学习算法的自动决策系统有很大的潜力对人们的生活产生不利影响。按照C部分的建议,将这些算法增强为隐私保护、公平和可解释的算法,可能会提高每个人的生活质量。该研究计划将支持13名受训人员,他们将了解机器学习方面的一系列最新进展,以成为下一代机器学习专家。作为一名女性初级调查员,我的目标是给女学生更多的实地性别平等培训机会。除此之外,我还打算包容代表性不足的群体,例如,在地理上(如非洲、南美洲和东南亚),或在种族上(土著人民)。
英文摘要
Data has great potential to provide world-changing solutions to pressing problems such as global hunger and climate change. However, sharing data at the scale required to tackle large problems comes with privacy concerns. My long-term goal is to provide methodological foundations that allow us to properly use and share these potentially instrumental data without privacy violations. Recent progress in privacy-preserving machine learning has centered around differential privacy, due to its provability. However, preserving privacy requires injecting noise in the machine learning system, leading to a fundamental trade-off between privacy and accuracy. The trade-off worsens when the algorithm's output is "sensitive" to any changes in its input, as the noise magnitude has to increase to hide that change. In Part A of the research program, I propose novel methods that can effectively limit the sensitivity. The second challenge arises due to the composability of differential privacy: every access to data reduces the privacy guarantee. Differentially private data generation solves this problem by creating a synthetic dataset that is reusable without limit. Unlike existing works assuming certain data distributions or particular downstream tasks, limiting the usefulness of the generated data, I propose a highly flexible nonparametric kernel-distance-based framework to compare the data and synthetic data distributions, in Part B. The third challenge arises due to the requirements for future algorithmic design imposed by regulations such as the General Data Protection Regulation. In Part C, I propose new frameworks that satisfy both differential privacy and other emerging notions, interpretability, causality, and fairness. Our theoretical and quantitative understanding of their interplay will be instrumental in developing practical algorithms that consider these notions simultaneously. There are two important societal impacts of the research program. First, the techniques like those in Part B ensure privacy protection which can promote more data sharing for the public good. Second, as the Netflix documentary Coded bias points out, automatic decision-making systems based on machine learning algorithms have a great potential to adversely affect people's lives. Augmenting those algorithms to be privacy-preserving, fair, and interpretable as suggested in Part C could potentially improve the quality of everyone's lives. The research program will support 13 trainees, who will learn about a broad range of recent advances in machine learning to be the next generation of machine learning experts. As a female primary investigator, my goal is to give more training opportunities to female students for gender equality in the field. Beyond that, I also intend to be inclusive to under-represented groups, e.g., geographically (such as Africa, South America, and Southeast Asia), or ethnically (indigenous people).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Turning an enemy into an ally: Privacy In Machine Learning (Pri-ML)
-
批准号:DGECR-2022-00376
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2022
-
负责人:Park, MiJung
-
依托单位:
海外基金