SLES: A Theoretical Lens on Generative AI Safety: Near and Long Term
SLES: A Theoretical Lens on Generative AI Safety: Near and Long Term
批准号:
2331831
负责人:
Sitan Chen
金额:
$80.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-11-01 至 2026-10-31
中文摘要
像ChatGPT这样的生成性人工智能技术以其合成惊人连贯的文本、代码等的能力席卷了世界。这些系统在质量上不断提高,并日益形成社会和工业的不同方面,这是值得注意的,但外地在控制和确保这些系统的可靠性方面的熟练程度还没有完全跟上。众所周知,这些模型仍然倾向于自信地做出事实不正确但听起来令人信服的声明。即使他们原则上拥有防止这种情况发生所需的所有知识,这些模型在拼凑起来时仍然经常出错。随着这项技术进入医疗保健或政策决策等关键任务环境,避免此类故障模式至关重要。这项研究将开发数学上严格的人工智能部署方法,这些方法具有坚实的理论保证,即系统不会以这种方式偏离其预期的行为。该项目的发现将有助于建立可持续的检查和故障安全机制,以便生成性人工智能技术能够以符合人类利益的受控方式进行扩展。这项研究旨在解决生成性人工智能在安全方面的短期挑战,以及随着这些模型能力的增长而出现的新的、长期的挑战。对于前者,该项目将在生成模型中为真实性和非幻觉性建立数学参数。这包括检测模型何时做出事实断言的实例,校准这些断言的置信度分数,可靠地将这些断言归因于它们在训练数据中的来源,以及鼓励模型在面对足够不分布的输入时放弃生成。另一个目标是研究获取和编辑存储在生成模型中的知识的方法,以及基于工具从细粒度复杂性理论和计算熵概念中分离出这样做的基本障碍。为了长期的安全,该项目将研究将紧急停止功能集成到基于加密后门的人工智能系统中的可行性,以及实施基于零知识证明的“人工智能武器协议”,以公开证明其安全属性,同时保持这些系统的某些组件的私密性。这项研究还将使用组合博弈论和递归启发式的平均案例分析技术,严格测试现有的人工智能系统可扩展监督建议,如自然语言辩论和迭代放大。这项研究由国家科学基金会和开放慈善机构之间的合作伙伴关系支持。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Generative AI technologies like ChatGPT have taken the world by storm with their ability to synthesize strikingly coherent text, code, and more. The pace with which these systems continue to improve in quality and increasingly shape diverse facets of society and industry is remarkable, yet the field's proficiency in controlling and ensuring the reliability of these systems has not quite kept up. These models remain notoriously prone to confidently making factually incorrect yet convincing-sounding statements. Even when they in principle have all of the knowledge that they need to prevent this, the models often still stumble in putting the pieces together. As this technology makes its way into mission-critical contexts like healthcare or policy decisions, it is crucial to avoid such failure modes. This research will develop mathematically rigorous AI deployment methods that come with solid theoretical assurances that the systems will not stray from their intended behavior in this way. The findings of this project will be instrumental in establishing sustainable checks and fail safes so that generative AI technologies can scale in a controlled fashion that is aligned with human interests. The research aims to tackle a mixture of both near-term challenges in safety for generative AI as well as emerging, longer-term ones that will arise as these models grow in their capabilities. For the former, the project will establish mathematical parameters for factuality and non-hallucination in generative models. This encompasses detecting instances when models make factual assertions, calibrating confidence scores for these assertions, reliably attributing these assertions to their sources in the training data, and encouraging models to abstain from generation when faced with sufficiently out-of-distribution input. Another goal is investigating methodologies to elicit and edit knowledge stored in generative models, as well as isolating fundamental barriers to doing so based on tools from fine-grained complexity theory and computational notions of entropy. For safety in the longer-term, the project will examine the feasibility of integrating emergency stop functionality into AI systems based on cryptographic backdoors, as well as implementing "AI arms protocols" based on zero knowledge proofs to publicly certify their safety properties while keeping certain components of these systems private. The research will also rigorously stress-test existing proposals for scalable oversight of AI systems, like natural-language debate and iterated amplification, using techniques from combinatorial game theory and average-case analysis of recursive heuristics.This research is supported by a partnership between the National Science Foundation and Open Philanthropy.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
PostDoctoral Research Fellowship
-
批准号:2103300
-
项目类别:Fellowship Award
-
资助金额:$15.0万
-
财政年份:2021
-
负责人:Sitan Chen
-
依托单位:
海外基金