课题基金 / 基金详情

Robustness and Interpretability of Foundation Models

Robustness and Interpretability of Foundation Models
基础模型的稳健性和可解释性
批准号:
2722135
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
简要描述研究背景,包括潜在影响:大型语言模型和基础模型在各种应用中得到广泛采用,例如,编程、网络搜索、写书、协助医生、客户服务、平面设计等。这项研究解决了生成式人工智能开发和部署中的重要挑战,如鲁棒性、公平性、隐私性、可解释性和安全使用。目标包括(i)形式化和基准化对自回归模型的攻击,(ii)新的安全评估方法,(iii)对系统即时泄漏的攻击,(iv)认证专业领域,(v)检查跨语言的公平性,(vi)统一鲁棒性和隐私,(vii)认证校准和集合,(vii)解释安全/不安全行为,以及(viii)开发前缀和后缀调谐理论。研究方法的新奇:该研究将引入新的想法,如(i)预算对抗提示,(ii)聊天风格的攻击,(iii)基于分布差异的目标,(iv)为差分隐私推广集合级Lipschitzness,以及(v)校准错误感知训练。与EPSRC的战略和研究领域保持一致:人工智能任何参与的公司或合作者:无
英文摘要
Brief description of the context of the research including potential impact: Large language models and foundation models are seeing widespread adoption in various applications, e.g., programming, web search, writing books, assisting medical practitioners, customer service, graphic design, etc. This research addresses important challenges in the development and deployment of generative AI, such as robustness, fairness, privacy, interpretability, and safe use.Aims and Objectives: Objectives include (i) formalizing and benchmarking attacks on autoregressive models, (ii) new evaluation methods for safety, (iii) attacks for system prompt leakage, (iv) certifying domains of expertise, (v) examining fairness across languages, (vi) unifying robustness and privacy, (vii) certifying calibration and ensembles, (vii) interpreting safe/unsafe behaviour, and (viii) developing a theory of prefix and suffix tuning.Novelty of the research methodology: The research would introduce novel ideas such as (i) budgeted adversarial prompts, (ii) chat-style attacks, (iii) distribution divergence-based objectives, (iv) generalizing set-level Lipschitzness for differential privacy, and (v) calibration error-aware training.Alignment to EPSRC's strategies and research areas: Artificial intelligenceAny companies or collaborators involved: None
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金