SaTC: CORE: Small: Decentralized Attribution and Secure Training of Generative Models
SaTC: CORE: Small: Decentralized Attribution and Secure Training of Generative Models
批准号:
2101052
负责人:
Yi Ren
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-10-01 至 2024-09-30
中文摘要
产生式模型描述了真实世界的数据分布,如图像、文本和人体运动,并在从照片编辑到自然语言处理再到自动驾驶的广泛且不断增长的应用中发挥着重要作用。关于生成性模型的开发和传播存在两个开放的挑战:(1)生成性模型的敌意应用已经针对社会技术干扰(例如,间谍活动和恶意冒充)而创建;(2)使用多个专有数据集(这是减少数据偏差所必需的)开发生成性模型会引起对数据泄露的隐私担忧。最近在这些挑战之后进行了立法努力,迄今为止,对条例的格式以及对其技术或社会可行性的了解有限。为此,该项目将开发新的数学理论和计算工具,以评估针对这些挑战的两种相互关联的解决方案的可行性:模型归属使所有者能够根据其生成的内容被正确识别;安全培训确保在可归因性生成模型的协作培训期间没有数据泄露。如果成功,该项目的成果将为今后的规章设计提供技术指导,以确保可再生模式的开发和传播。项目成果将通过项目网站、开放源码软件和公共数据集传播。该项目的影响将通过教育活动扩大,包括关于人工智能(AI)安全的新课程模块、本科研究项目,以及通过实验室参观延伸到当地社区,以培养具有技能的代表不足的群体,以减轻针对这些群体的恶意模拟和有偏见的数据/模型表示的风险。该项目将侧重于协同研究任务,以实现分散的模型归属和安全的生成模型培训。在前者中,研究团队将研究一套用户端生成性模型的系统设计,这些模型可以由一组二进制分类器来证明属性,这些二进制分类器以分散的方式存储以降低安全风险。分散归属的技术可行性将通过在可归属性、发电质量和模型容量之间的权衡来衡量。在后者中,研究小组将研究生成模型和相关的二进制分类器的安全多方训练以进行归因。将通过设计安全友好的模型架构和学习损失来平衡数据隐私和培训的可扩展性。将创造新的知识,将本项目与数字取证和安全计算方面的现有最先进文献区分开来:(1)将开发分散归属的充分条件,这将揭示可归因性、数据几何、模型架构和生成质量之间的分析联系。(2)充分的条件将能够估计给定数据集的属性模型的容量和生成质量容差。(3)研究了次线性安全向量乘法的可行性,从根本上提高了安全协同训练的可扩展性。(4)隐私友好的激活和丢失功能将用于培训用户端生成模型和用于归因的分类器。该奖项反映了NSF的法定使命,并已通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Generative models describe real-world data distributions such as images, texts, and human motions, and are playing an essential role in a large and growing range of applications from photo editing to natural language processing to autonomous driving. There are two open challenges regarding the development and dissemination of generative models: (1) Adversarial applications of generative models have created concerning socio-technical disturbances (e.g., espionage operations and malicious impersonation); and (2) developing generative models using multiple proprietary datasets (which are needed to reduce data biases) raises privacy concerns about data leakage. Legislative efforts have recently been taken in the wake of these challenges, so far with limited consensus on the format of regulations and knowledge about their technological or social feasibility. To this end, this project will develop new mathematical theories and computational tools to assess the feasibility of two connected solutions to these challenges: Model attribution enforces the owners to be correctly identified based on their generated contents; secure training ensures zero data leakage during the collaborative training of attributable generative models. If successful, the outcomes of the project will provide technical guidance for future regulation design towards secure development and dissemination of generative models. Project results will be disseminated through a project website, open-source software, and public datasets. The impacts of the project will be broadened through educational activities, including new course modules on Artificial Intelligence (AI) security, undergraduate research projects, and outreach to the local community through lab tours, to prepare underrepresented groups with skills to mitigate risks from malicious impersonation and biased data/model representations targeting these groups.This project will focus on synergistic research tasks towards decentralized model attribution and secure training of generative models. In the former, the research team will study the systematic design of a set of user-end generative models that can be certifiably attributed by a set of binary classifiers, which are stored in a decentralized manner to mitigate security risks. The technical feasibility of decentralized attribution will be measured by the tradeoffs between attributability, generation quality, and model capacity. In the latter, the research team will study secure multi-party training of generative models and the associated binary classifiers for attribution. Data privacy and training scalability will be balanced through the design of security-friendly model architectures and learning losses. New knowledge will be created that differentiates this project from the existing state-of-the-art literature in digital forensics and secure computation: (1) Sufficient conditions for decentralized attribution will be developed, which will reveal analytical connections between attributability, data geometry, model architecture, and generation quality. (2) The sufficient conditions will enable estimation of the capacity of attributable models for a given dataset and generation quality tolerance. (3) Feasibility of sublinear secure vector multiplication will be studied, which will fundamentally improve the scalability of secure collaborative training. (4) Privacy-friendly activation and loss functions will be designed for the training of user-end generative models and the classifiers for attribution.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
作者:
[Changhoon Kim;Yi Ren;Yezhou Yang]
通讯作者:
Changhoon Kim;Yi Ren;Yezhou Yang
DOI:
10.48550/arxiv.2304.09752
发表时间:
2023-04
期刊:
影响因子:
--
作者:
[Guangyu Nie;C. Kim;Yezhou Yang;Yi Ren]
通讯作者:
Guangyu Nie;C. Kim;Yezhou Yang;Yi Ren
DOI:
10.1109/icassp43922.2022.9746578
发表时间:
2022
期刊:
Proceedings of the IEEE International Conference on Acoustics Speech and Signal Processing
影响因子:
--
作者:
[Cho, Yongbaek, Kim, Changhoon, Yang, Yezhou, Ren, Yi]
通讯作者:
Ren, Yi
DOI:
10.1145/3460120.3484778
发表时间:
2021-11
期刊:
Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
作者:
[Mike Rosulek;Ni Trieu]
通讯作者:
Mike Rosulek;Ni Trieu
DOI:
10.1007/978-3-030-92075-3_21
发表时间:
2020
期刊:
IACR Cryptol. ePrint Arch.
影响因子:
--
作者:
[Tancrède Lepoint;Sarvar Patel;Mariana Raykova;Karn Seth;Ni Trieu]
通讯作者:
Tancrède Lepoint;Sarvar Patel;Mariana Raykova;Karn Seth;Ni Trieu
DMS/NIGMS 2: Collaborative Research: Developing Statistical Learning Methods for Revealing the Molecular Signatures of Microvascular Changes in Neural Injury
-
批准号:2054014
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2021
-
负责人:Yi Ren
-
依托单位:
Collaborative Research: Statistical Methods for RNA-seq Based Transcriptomic Analysis of Macrophage Function in Spinal Cord Injury
-
批准号:1661727
-
项目类别:Continuing Grant
-
资助金额:$80.0万
-
财政年份:2017
-
负责人:Yi Ren
-
依托单位:
EAGER: Reconstruction and Optimal Design of Multi-scale Material Systems through Deep Networks
-
批准号:1651147
-
项目类别:Standard Grant
-
资助金额:$17.13万
-
财政年份:2016
-
负责人:Yi Ren
-
依托单位:
Collaborative Research: Development of bioinformatic methods for studying gene expression network inflammation and neuronal regeneration
-
批准号:1419553
-
项目类别:Continuing Grant
-
资助金额:$28.37万
-
财政年份:2013
-
负责人:Yi Ren
-
依托单位:
Collaborative Research: Development of bioinformatic methods for studying gene expression network inflammation and neuronal regeneration
-
批准号:0714589
-
项目类别:Continuing Grant
-
资助金额:$81.0万
-
财政年份:2007
-
负责人:Yi Ren
-
依托单位:
国内基金
海外基金
登录
查看更多内容
胆固醇羟化酶CH25H非酶活依赖性促进乙型肝炎病毒蛋白Core及Pre-core降解的分子机制研究
-
批准号:82371765
-
项目类别:面上项目
-
资助金额:50万元
-
批准年份:2023
-
负责人:谭广云
-
依托单位:
锕系元素5f-in-core的GTH赝势和基组的开发
-
批准号:22303037
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:鲁俊波
-
依托单位:
基于合成致死策略搭建Core-matched前药共组装体克服肿瘤耐药的机制研究
-
批准号:--
-
项目类别:--
-
资助金额:52万元
-
批准年份:2022
-
负责人:孙丙军
-
依托单位:
鼠伤寒沙门氏菌LPS core经由CD209/SphK1促进树突状细胞迁移加重炎症性肠病的机制研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:叶成林
-
依托单位:
基于外泌体精准调控的“核-壳”(core-shell)同步血管化骨组织工程策略的应用与机制探讨
-
批准号:--
-
项目类别:--
-
资助金额:55万元
-
批准年份:2020
-
负责人:张智勇
-
依托单位:
基于外泌体精准调控的“核-壳”(core-shell)同步血管化骨组织工程策略的应用与机制探讨
-
批准号:82072415
-
项目类别:面上项目
-
资助金额:55.0万元
-
批准年份:2020
-
负责人:张智勇
-
依托单位:
肌营养不良蛋白聚糖Core M3型甘露糖肽的精确制备及功能探索
-
批准号:92053110
-
项目类别:重大研究计划
-
资助金额:70.0万元
-
批准年份:2020
-
负责人:彭鹏
-
依托单位:
Core-1-O型聚糖黏蛋白缺陷诱导胃炎发生并介导慢性胃炎向胃癌转化的分子机制研究
-
批准号:81902805
-
项目类别:青年科学基金项目
-
资助金额:20.5万元
-
批准年份:2019
-
负责人:刘菲
-
依托单位:
原始地球增生晚期的Core-merging大碰撞事件:地核增生、核幔平衡与核幔边界结构的新认识
-
批准号:41973063
-
项目类别:面上项目
-
资助金额:65.0万元
-
批准年份:2019
-
负责人:周游
-
依托单位:
CORDEX-CORE区域气候模拟与预估研讨会
-
批准号:41981240365
-
项目类别:国际(地区)合作与交流项目
-
资助金额:1.5万元
-
批准年份:2019
-
负责人:陈威霖
-
依托单位: