CAREER: Towards Provenance-Driven Understanding of Machine Learning Robustness
CAREER: Towards Provenance-Driven Understanding of Machine Learning Robustness
批准号:
2238084
负责人:
Birhanu Eshete
金额:
$61.98万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-05-01 至 2028-04-30
中文摘要
机器学习(ML)越来越多地用于社会关键应用,如自动驾驶汽车、医学、金融和刑事司法。然而,ML也容易受到攻击者的攻击,他们可以攻击数据模型和ML模型本身。这可能导致模型中的不良行为和使用它们的人的不良决策。该项目的目标是通过专注于来源来提高我们检测和响应攻击的能力:系统地捕获数据和构建模型时使用的训练方法,沿着部署后的推理过程和决策。通过捕获这些数据并开发在评估风险、审计模型和取证分析事件时使用这些数据的方法,这项工作将使ML系统在攻击方面更加强大和负责。这些功能反过来将使开发和使用ML模型的组织以及监督其效果的政策制定者和监管机构沿着受益。这项工作分为三个主要方面。第一个重点是系统地捕获和表征部署前(训练)和部署后(推理)的起源,重点是什么构成训练和推理起源以及ML计算的固有非确定性。特别是,训练和推理元数据,训练进展,推理计算动态和每个标签的表征方法将被探索。第二个重点将使用这些数据进行来源驱动的检测,以检测一系列威胁模型和应用程序域中的训练数据中毒和模型规避。对于中毒检测,基于相似性和基于分布偏移检测的方法将被追求,而对于逃避检测,推理出处将被经验和结构化地分析。第三个重点是发展妥协后的取证能力,目标是追溯攻击的原因并减轻未来的攻击。与这三个目标相结合的是一个教育计划,包括为本科生和研究生开发关于ML可信度的新课程,为本科生举办强大的ML道德黑客竞赛,而K-12个关于强大ML的夏令营,旨在培养下一代网络安全工作者并使其多样化。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准。
英文摘要
Machine Learning (ML) is increasingly used in socially critical applications such as self-driving cars, medicine, finance, and criminal justice. However, ML is also susceptible to adversaries who can attack both the data models are trained on and the ML models themselves. This can lead to poor behavior in the models and poor decisions in the people who use them. This project’s goal is to advance our ability to detect and respond to attacks through focusing on provenance: systematic capture of the data and training methods used in building models, along with the inference processes and decisions made after they are deployed. By capturing these data and developing methods to use the data when assessing risks, auditing models, and forensically analyzing incidents, the work will make ML systems both more robust and more accountable around attacks. These capabilities will in turn benefit organizations that develop and use ML models, along with policymakers and regulators who oversee their effects. The work is organized into three main thrusts. The first thrust focuses on systematic capture and characterization of pre-deployment (training) and post-deployment (inference) provenance, focusing on what constitutes training and inference provenance and the innate nondeterminism of ML computations. In particular, training and inference metadata, training progression, inference computation dynamics, and per-label characterization approaches will be explored. The second thrust will use these data for provenance-driven detection of training data poisoning and model evasion across a range of threat models and application domains. For poisoning detection, both similarity-based and distribution shift detection-based approaches will be pursued while for evasion detection, inference provenance will be analyzed empirically and structurally. The third thrust focuses on developing post-compromise forensics capabilities with the goal of tracing back attacks to their cause(s) and mitigating future attacks. Integrated with these three thrusts is an educational plan that includes developing new courses on ML trustworthiness for undergraduate and graduate students, robust ML-focused ethical hacking competitions for undergraduates, and K-12 summer camps on robust ML to develop and diversify the next generation of cybersecurity workers.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DeResistor: Toward Detection-Resistant Probing for Evasion of Internet Censorship
DeResistor:针对逃避互联网审查的抗检测探测
DOI:
--
发表时间:
2023
期刊:
32nd USENIX Security Symposium
影响因子:
--
作者:
[Amich, Abderrahmen, Eshete, Birhanu, Yegneswaran, Vinod, Hoang, Nguyen Phong]
通讯作者:
Hoang, Nguyen Phong
海外基金