Robustness and Interpretability of Foundation Models
Robustness and Interpretability of Foundation Models
批准号:
2722135
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Brief description of the context of the research including potential impact: Large language models and foundation models are seeing widespread adoption in various applications, e.g., programming, web search, writing books, assisting medical practitioners, customer service, graphic design, etc. This research addresses important challenges in the development and deployment of generative AI, such as robustness, fairness, privacy, interpretability, and safe use.Aims and Objectives: Objectives include (i) formalizing and benchmarking attacks on autoregressive models, (ii) new evaluation methods for safety, (iii) attacks for system prompt leakage, (iv) certifying domains of expertise, (v) examining fairness across languages, (vi) unifying robustness and privacy, (vii) certifying calibration and ensembles, (vii) interpreting safe/unsafe behaviour, and (viii) developing a theory of prefix and suffix tuning.Novelty of the research methodology: The research would introduce novel ideas such as (i) budgeted adversarial prompts, (ii) chat-style attacks, (iii) distribution divergence-based objectives, (iv) generalizing set-level Lipschitzness for differential privacy, and (v) calibration error-aware training.Alignment to EPSRC's strategies and research areas: Artificial intelligenceAny companies or collaborators involved: None
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金