CIF: Medium: Collaborative Research: Information-theoretic Guarantees on Privacy in the Age of Learning
CIF: Medium: Collaborative Research: Information-theoretic Guarantees on Privacy in the Age of Learning
批准号:
1901243
负责人:
Lalitha Sankar
金额:
$81.7万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
未结题
起止时间:
2019-06-01 至 2025-05-31
中文摘要
随着机器学习的强大进步,感兴趣的一方从个人不断扩大的数字足迹中收集个人信息的能力正在超过任何人保持信息隐私的能力。虽然这些聚合数据可以通过建立在机器学习和人工智能基础上的技术为消费者和数据科学家带来巨大好处,但这种好处必须得到有意义的隐私保证,而这些隐私正是最初提供数据的人。为此,该项目探索了在学习环境中隐私泄露的有意义的测量方法,表征了隐私和效用之间的基本权衡,开发了在现实环境中确保隐私的技术,并在公开可用的数据集上测试了这些算法。该项目还致力于通过两个外展努力扩大对计算机的参与:(I)通过亚利桑那州立大学一年一度的STEM活动Open Door向初中生展示源于使用社交媒体的隐私问题的互动演示,以及(Ii)亚利桑那州立大学通过青年工程师塑造世界(Yesw)暑期计划和哈佛大学关于机器学习(ML)和人工智能(AI)的教学模块和短期课程(数据堵塞);外展工作将通过亚利桑那州立大学S学院研究和评估服务团队(CREST)使用广为人知的评估学生兴趣、参与度和知识的指标进行评估。该项目旨在得出一个基本的、统计的隐私理论,该理论建立在信息理论和机器学习的现代理论进步的基础上,并对其做出贡献。从这一观点中衍生出的一个重要的新元素是最大α泄漏,这是一种新的、可调的信息泄漏度量,它量化了对手通过损失函数的参数类学习私人数据的任何功能的能力。这一可调整的衡量标准源于基于Renyi分歧的丰富信息理论框架,从而将不同的现有衡量标准统一在一个框架下。此外,它的操作意义和计算灵活性允许在机器学习中自然应用。在这些措施的背景下,本项目从理论上和以数据驱动的方式在两个不同的背景下研究隐私-效用权衡:(I)以与原始数据相似的形式发布数据集,并为任意统计分析提供隐私和严格的效用保证;(Ii)为特定的学习任务发布隐私保证的数据表示。
英文摘要
Armed with powerful advances in machine learning, the ability of an interested party to gather personal information from an individual's expanding digital footprint is outstripping anyone's capability to keep their information private. While this aggregated data can have tremendous benefit for consumers and data scientists via technologies built on machine learning and artificial intelligence, this benefit must be tempered with meaningful assurances of privacy for the very people who provided the data in the first place. This project adopts a rigorous information-theoretic approach to give meaningful privacy guarantees while still providing statistical utility. By combining theoretical and data-driven research, this project can inform public policy as well as best-practices for industry. The overall goal is to provide any data scientist with a set of tools to guarantee meaningful privacy in practice. To do so, this project explores meaningful measures of privacy leakage in the learning context, characterizes the fundamental tradeoffs between privacy and utility, develops techniques to ensure privacy in realistic settings, and tests these algorithms on publicly available datasets. The project is also committed to broadening participation in computing via two outreach efforts: (i) interactive demonstrations of privacy issues that stem from using social media to middle and high school students via ASU's annual STEM event, Open Door, and (ii) teaching modules on machine learning (ML) and artificial intelligence (AI), and short courses ("data jams") at ASU via the Young Engineers Shape the World (YESW) summer program and at Harvard; these modules, targeted at female, financially disadvantaged, and Latino and Hispanic students, aim to make a meaningful contribution to increasing a diverse STEM workforce by providing students hands-on experience on basic concepts of coding, manipulating datasets, and producing simple visualizations collectively. Outreach efforts will be evaluated using well understood metrics for assessment of student interest, engagement, and knowledge via ASU?s College Research and Evaluation Services Team (CREST).This project aims to derive a foundational, statistical theory of privacy that builds upon and contributes to modern theoretical advances in information theory and machine learning. The statistical nature of inference (both for legitimate and illegitimate ends) requires a statistical approach to measuring and ensuring privacy and utility. A significant novel element derived from this view is the maximal alpha leakage, a new, tunable measure for information leakage which quantifies the ability of an adversary to learn any function of private data via a parametric class of loss functions. This tunable measure is derived from a rich information-theoretic framework based on Renyi divergence, thereby uniting disparate existing measures under a single framework. Moreover, its operational significance and computational flexibility allow for natural application in machine learning. In the context of these measures, this project studies privacy-utility tradeoffs both theoretically and in a data-driven manner in two distinct settings: (i) releasing datasets in a similar form as the original, with privacy and strict utility guarantees for arbitrary statistical analysis, and (ii) releasing privacy-guaranteed data representations for specific learning tasks. Broader dissemination of the work will go beyond conferences to organizing a privacy workshop in the latter half of the project to enable inter-disciplinary interactions and application.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(22)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/tit.2019.2935768
发表时间:
2019-12-01
期刊:
IEEE TRANSACTIONS ON INFORMATION THEORY
影响因子:
2.5
作者:
[Liao, Jiachun, Kosut, Oliver, Calmon, Flavio du Pin]
通讯作者:
Calmon, Flavio du Pin
Evaluating Multiple Guesses by an Adversary via a Tunable Loss Function
通过可调谐损失函数评估对手的多个猜测
DOI:
10.1109/isit45174.2021.9517733
发表时间:
2021
期刊:
International Symposium on Information Theory
影响因子:
--
作者:
[Kurri, Gowtham R., Kosut, Oliver, Sankar, Lalitha]
通讯作者:
Sankar, Lalitha
DOI:
--
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
作者:
[R. Nock;Tyler Sypherd;L. Sankar]
通讯作者:
R. Nock;Tyler Sypherd;L. Sankar
α-GAN: Convergence and Estimation Guarantees
α-GAN:收敛和估计保证
DOI:
10.1109/isit50566.2022.9834890
发表时间:
2022
期刊:
IEEE International Symposium on Information Theory
影响因子:
--
作者:
[Kurri, Gowtham R., Welfert, Monica, Sypherd, Tyler, Sankar, Lalitha]
通讯作者:
Sankar, Lalitha
DOI:
10.1109/tit.2019.2939472
发表时间:
2020
期刊:
IEEE Transactions on Information Theory
影响因子:
2.5
作者:
[Diaz, Mario, Wang, Hao, Calmon, Flavio P., Sankar, Lalitha]
通讯作者:
Sankar, Lalitha
共 17 条
Exploiting Physical and Dynamical Structures for Real-time Inference in Electric Power Systems
-
批准号:2246658
-
项目类别:Standard Grant
-
资助金额:$36.0万
-
财政年份:2023
-
负责人:Lalitha Sankar
-
依托单位:
Collaborative Research: SCH: Fair Federated Representation Learning for Breast Cancer Risk Scoring
-
批准号:2205080
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2022
-
负责人:Lalitha Sankar
-
依托单位:
Unifying Information- and Optimization-Theoretic Approaches for Modeling and Training Generative Adversarial Networks
-
批准号:2134256
-
项目类别:Continuing Grant
-
资助金额:$110.0万
-
财政年份:2021
-
负责人:Lalitha Sankar
-
依托单位:
RAPID: SaTC: FACT: Federated Analytics based Contact Tracing for COVID-19
-
批准号:2031799
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2020
-
负责人:Lalitha Sankar
-
依托单位:
CIF: Small: Alpha Loss: A New Framework for Understanding and Trading Off Computation, Accuracy, and Robustness in Machine Learning
-
批准号:2007688
-
项目类别:Standard Grant
-
资助金额:$50.8万
-
财政年份:2020
-
负责人:Lalitha Sankar
-
依托单位:
Student Travel Support for the 2020 IEEE SGComm Conference. To be Held November, 11-13, 2020 at Arizona State University.
-
批准号:2024805
-
项目类别:Standard Grant
-
资助金额:$0.88万
-
财政年份:2020
-
负责人:Lalitha Sankar
-
依托单位:
Collaborative Research: High-Dimensional Spatio-Temporal Data Science for a Resilient Power Grid: Towards Real-Time Integration of Synchrophasor Data
-
批准号:1934766
-
项目类别:Continuing Grant
-
资助金额:$131.4万
-
财政年份:2019
-
负责人:Lalitha Sankar
-
依托单位:
CIF: Small: Collaborative Research: Generative Adversarial Privacy: A Data-driven Approach to Guaranteeing Privacy and Utility
-
批准号:1815361
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2018
-
负责人:Lalitha Sankar
-
依托单位:
CPS: TTP Option: Synergy: A Verifiable Framework for Cyber- Physical Attacks and Countermeasures in a Resilient Electric Power Grid
-
批准号:1449080
-
项目类别:Cooperative Agreement
-
资助金额:$140.0万
-
财政年份:2015
-
负责人:Lalitha Sankar
-
依托单位:
CAREER: Privacy-Guaranteed Distributed Interactions in Critical Infrastructure Networks
-
批准号:1350914
-
项目类别:Continuing Grant
-
资助金额:$45.5万
-
财政年份:2014
-
负责人:Lalitha Sankar
-
依托单位:
海外基金