课题基金 / 基金详情

Security and compilers for machine learning

Security and compilers for machine learning
机器学习的安全性和编译器
批准号:
2906291
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2024
资助国家:
英国
项目状态:
未结题
起止时间:
2024 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
机器学习(ML)正在各个行业中迅速获得吸引力,有望在各个领域带来变革性的好处。然而,对ML系统的日益依赖揭示了对强大的安全和安全措施的迫切需求。这是由于与ML模型相关的固有漏洞及其滥用的潜在后果。其中一个主要问题是ML模型对对抗性攻击的敏感性,其中恶意行为者操纵数据,模型参数或模型架构以利用系统。这些攻击可能导致偏见,不准确,甚至危险的决策。此外,机器学习模型的复杂性使得识别和缓解漏洞变得非常具有挑战性,这使得它们难以防御。另一个重要问题是AI对齐。对齐是指确保人工智能(AI)系统的行为方式与人类价值观和目标保持一致的过程。它涉及开发技术来指导人工智能模型做出有利于人类的决策和行动,同时最大限度地减少潜在的危害。AI对齐对于AI系统的负责任开发和部署至关重要,因为它有助于确保AI技术与人类利益保持一致,并以道德和有益的方式使用。该博士探讨了提高ML安全性和安全性的各种途径。最初的项目如下:一个项目是提高用于微调ML模型的人类偏好数据的质量。校准在很大程度上依赖于人类偏好的质量,但现有的大部分数据都是由过度劳累和收入不足的工人产生的,他们没有真实的动机来提供好的数据。该项目将实验性地研究根据员工的偏好是否成功地提高了模型在现有基准上的性能来支付员工奖金的效果。该研究还将确定以这种方式增加人类动机是否会提高模型的性能,即使在没有奖励的指标中也是如此。如果是这样的话,这可能会导致在没有好的基准的指标上有更好的表现,例如政治偏见。另一个提高模型对抗性攻击安全性的项目是调查对自我奖励语言模型的新推动是否为后门创造了机会,通过这些自我奖励模型中固有的反馈回路被放大。这源于模型崩溃等想法,其中对LLM生成的数据进行训练可能导致总体性能失败,以及现有的数据中毒工作,以便在LLM中插入后门。第三个项目是研究将机器学习模型锁定到特定硬件的各种方法,例如通过使用难以伪造的硬件指纹(例如基于完成操作所需的时钟周期的数量)作为用于模型的权重的加密密钥,或通过优化模型的特定量化方案,只存在于一些硬件。这个项目符合EPSRC研究领域“人工智能技术”。
英文摘要
Machine learning (ML) is rapidly gaining traction across various industries, promising transformative benefits in diverse fields. However, the increasing reliance on ML systems has brought to light the crucial need for robust security and safety measures. This is due to the inherent vulnerabilities associated with ML models and the potential consequences of their misuse.One of the primary concerns is the susceptibility of ML models to adversarial attacks, where malicious actors manipulate data, model parameters, or model architecture to exploit the system. These attacks can result in biased, inaccurate, and even dangerous decision-making. Additionally, the complexity of ML models makes it challenging to identify and mitigate vulnerabilities, making them difficult to defend against.Another significant issue is AI alignment. Alignment refers to the process of ensuring that artificial intelligence (AI) systems behave in ways that align with human values and objectives. It involves developing techniques to guide AI models towards making decisions and taking actions that are beneficial to humanity, while minimizing potential harms. AI alignment is crucial for the responsible development and deployment of AI systems, as it helps ensure that AI technologies align with human interests and are used ethically and beneficially.This PhD explores various avenues in improving ML security and safety. The initial projects are as follows.One project is to improve the quality of human preference data used to fine-tune ML models. Alignment relies heavily on the quality of human preferences, but much of the existing data is generated by overworked and underpaid workers with no real incentive to provide good data. This project would experimentally research the effect of paying workers bonuses based on whether their preferences successfully improve the performance of the model on existing benchmarks. The research will also determine whether increasing human motivation in this way increases the performance of the model even in metrics which are not rewarded. If so, this could lead to better performance in metrics for which there are no good benchmarks, such as political bias.Another project to improve the security of models against adversarial attack is to investigate whether the new push towards self-rewarding language models creates an opportunity for backdoors to be amplified through the inherent feedback loop in these self-rewarding models. This follows from ideas such as Model Collapse, in which training on LLM-generated data can lead to total performance failure, and the existing body of work on data poisoning to insert backdoors in LLMs.A third project is to investigate various methods for locking machine learning models to specific hardware, such as by using a difficult-to-forge hardware fingerprint (e.g. based on the number of clock cycles required to complete an operation) as an encryption key for the weights of the model, or by optimising models for particular quantisation schemes that only exist on some hardware.This project aligns with the EPSRC research area "Artificial intelligence technologies".
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金