课题基金 / 基金详情

Towards Trustworthy Large Language Models

Towards Trustworthy Large Language Models
迈向可信赖的大型语言模型
批准号:
2895111
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在过去的几年中,大型语言模型(广义上讲是基础模型)(例如ChatGPT、GPT-3 Brown等人)。[2020],GPT-4 OpenAI[2023])掀起了人工智能(AI)领域的热潮。更具体地说,随着最近于2022年11月发布的ChatGPT,更多的观众体验到了LLMS的生成能力。大语言模型的生成能力已经成功地应用于自然语言处理任务的不同领域。伴随着革命性的影响,人们对在不同应用中使用LLMS的利害关系提出了许多问题。总的来说,科学界的很大一部分人建议以一种对社会负责和符合伦理的方式使用LLMS,NAT[2023]。因此,这个项目的目标是建立可解释的低成本模型。LLMS的最终用户可以是不同类型的。用户可以是使用在其后端使用LLMS的NLP模型的领域专家,或者是投资于使用LLMS的AI产品的利益相关者,或者是没有AI专业知识的人。每种类型的用户都应该能够信任LLMS提供的输出。现有研究表明,向用户解释人工智能模型的输出应该有助于增加用户对系统的信任。广义地说,可解释性的概念是通过一个简单的解释器模块来理解AI模型的工作原理,该模块可以模仿原始的AI模型。在这个项目中,我们将特别侧重于向每种类型的用户(即领域专家、利益相关者、普通人)解释LLMS的输出。这项研究建议的总体目标是使用可解释性技术增加低成本管理的透明度。除了透明度,可解释的LLM还可以帮助识别模型本身中存在的任何类型的偏见。最终可解释的LLMS是朝着创建一个对社会负责的人工智能环境的目标迈出的一步。
英文摘要
In the past few years Large Language models (broadly speaking foundational models) (e.g. ChatGPT, GPT-3 Brown et al. [2020], GPT-4 OpenAI [2023]) have stirred up the field of Artificial Intelligence (AI). More specifically, with the recent release of ChatGPT in November 2022, a wider section of audience got to experience the generative power of LLMs. The generative power of large language models (LLM) has been successfully applied in different areas of natural language processing tasks. Along with the revolutionary impact, many questions have been raised regarding the stakes of using LLMs in different applications. Broadly speaking a significant portion of the scientific community has advised to use LLMs in a socially responsible and ethical way Nat [2023]. Consequently, the aim of this project is to build explainable LLMs. The end user for LLMs can be of different types. The user may be a domain expert using an NLP model which uses LLMs at its back end or a stakeholder, investing in an AI product, which uses LLMs or someone having no AI expertise. Each type of user should be able to trust the output provided by LLMs. Existing research has shown that explaining the output of an AI model to a user should help to increase a user's trust in the system. Broadly speaking, the idea of explainability is to understand the working principle of an AI model with a simple explainer module which can mimic the original AI model. In this project we would like to specifically focus on explaining the output of LLMs to every type of users (i.e. domain experts, stakeholders, common people). The overall goal of this research proposal is to increase transparency of the LLMs using explainability techniques. Along with transparency, explainable LLM can also help to identify any kind of bias present in the model itself. Eventually explainable LLMs is a step towards the goal of creating a socially responsible AI environment.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金