课题基金 / 基金详情

Collaborative Research: SaTC: CORE: Small: Differentially Private Data Synthesis: Practical Algorithms and Statistical Foundations

Collaborative Research: SaTC: CORE: Small: Differentially Private Data Synthesis: Practical Algorithms and Statistical Foundations
协作研究:SaTC:核心:小型:差分隐私数据合成:实用算法和统计基础
批准号:
2247794
负责人:
Ninghui Li
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-15 至 2026-06-30

项目摘要

项目成果

Ninghui Li的其他基金

相似基金

相关文献

中文摘要
翻译
组织和机构收集的数据是当今信息时代的关键资源,是当今经济的重要组成部分。然而,这些数据的泄露对个人隐私构成了严重威胁。在保护隐私的同时使用数据的一个重要方法是差分私有数据合成(DPDS)。也就是说,给定私有数据集作为输入,可以使用差分私有算法来生成与输入数据集“相似”的合成数据集。虽然DPDS近年来受到了极大的关注,但我们对这一主题的理解仍然有限。这个项目采取了多学科的方法,以增进我们对DPDS的科学理解,并改进实践技术。更具体地说,这个项目的新颖性如下。首先,它系统地探索了基于边际的DPDS算法的设计空间,这些算法在DPDS上的NIST竞争中被证明是有效的,同时也从类似领域(通常不能满足DP)开发的数据合成技术中获得见解。其次,发展了基于DPDS算法经验性能的统计理论,并指导了这些算法的实证研究。该项目更广泛的意义和重要性如下。我们正处于信息经济时代。正在收集各种数据,如在线交互、医学传感器数据、基因组数据和位置数据。迫切需要能够在保护个人隐私的同时使用这些数据的实用技术,并将极大地提高这些数据的价值。用户将受益于加强对其私人信息的控制,整个社会将受益于从聚合数据中获得最大利益。PIS计划在这一领域的现有研究以及该项目的研究成果的基础上,共同开发和教授一门关于合成数据的研究生课程,并让本科生参与研究。这个项目有两个突破口。第一个推力旨在开发新的基于边际的DPDS算法,以改进经验评估中的最新技术。这些任务包括:深入研究“边缘到数据集”问题(在给定一组边缘时如何合成数据集);开发和评价处理数字属性的新方法;开发选择边缘的自适应和自动化技术,以便用它们合成的数据集尽可能多地从输入数据集中获取有用的信息。第二个推力是对第一个推力的实证研究的补充,旨在发展基于高维边际的数据合成算法的统计理论,并建立一个通用的学习理论框架来评估合成数据在下游任务中的效用。这两个推进具有很强的互补性和互补性。推力1中的实验研究将为推力2中的理论研究提供见解和方向,这将有助于解释实验结果并指导其他实验研究。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Data collected by organizations and agencies are a key resource in today’s information age and fuel a significant part of today's economy. However, the disclosure of those data poses serious threats to individual privacy. One important approach to using data while protecting privacy is differential private data synthesis (DPDS). That is, given as input a private dataset, one uses a differentially private algorithm to generate synthetic datasets that are “similar” to the input dataset. While DPDS has received much attention in recent years, our understanding on this topic remains limited. This project takes a multi-disciplinary approach to advance our scientific understanding as well as improve practice techniques for DPDS. More specifically, this project’s novelties are as follows. First, it systematically explores the design space in marginal-based DPDS algorithms that have been proven to be effective in NIST competitions on DPDS, while also taking insights from data synthesis techniques developed in similar fields (often not satisfying DP). Second, it develops statistical theories that both are motivated by the empirical performances of DPDS algorithms, and guide the empirical research of these algorithms. The project’s broader significance and importance are as follows. We are in the information economy. Data of all kinds, such as online interaction, medical sensor data, genomic data, and location data are being collected. Practical techniques that enable use of these data while protecting individual privacy are crucially needed and will greatly enhance the value of such data. Users will gain from increased control of their private information, and society as a whole will benefit from deriving maximal benefit from aggregated data. PIs plan to jointly develop and teach a graduate-level course on synthetic data based on the existing research in this area as well as research results from this project, and involve undergraduate students in research. This project has two thrusts. The first thrust aims to develop new marginal-based DPDS algorithms that improve upon the state-of-art in empirical evaluations. The tasks include: perform an in-depth study of the “marginal-to-dataset” problem (how to synthesize a dataset when given a set of marginals); develop and evaluate new approaches for handling numerical attributes; and develop adaptive and automated techniques for selecting marginals so that dataset synthesized with them captures as much useful information from the input dataset as possible. The second thrust complements the empirical research in the first thrust, and aims to develop statistical theory for high dimensional marginal-based data synthesis algorithms, and also a general learning theory framework to evaluate the utility of synthetic data in downstream tasks. The two thrusts are highly complementary and support each other. The experimental study in Thrust 1 will provide insights and directions for theoretical studies in Thrust 2, which will help explain the experimental findings as well as guide additional experimental studies.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Proposal: SaTC: Frontiers: Center for Distributed Confidential Computing (CDCC)
  • 批准号:
    2207204
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $88.0万
  • 财政年份:
    2022
  • 负责人:
    Ninghui Li
  • 依托单位:
SaTC: CORE: Medium: Collaborative: User-Centered Deployment of Differential Privacy
  • 批准号:
    1931443
  • 项目类别:
    Standard Grant
  • 资助金额:
    $34.31万
  • 财政年份:
    2020
  • 负责人:
    Ninghui Li
  • 依托单位:
RAPID: Collaborative: PPSRC: Privacy-Preserving Self-Reporting for COVID-19
  • 批准号:
    2034235
  • 项目类别:
    Standard Grant
  • 资助金额:
    $13.31万
  • 财政年份:
    2020
  • 负责人:
    Ninghui Li
  • 依托单位:
SaTC: CORE: Improving Password Ecosystem: A Holistic Approach
  • 批准号:
    1704587
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2017
  • 负责人:
    Ninghui Li
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)