Benchmarked Ethics: A Roadmap to AI Alignment, Moral Knowledge, and Control

Benchmarked Ethics: A Roadmap to AI Alignment, Moral Knowledge, and Control
复制标题

道德基准:人工智能联盟、道德知识和控制的路线图

DOI:
10.1145/3600211.3604764
复制
发表时间:
2023
期刊:
and Society
影响因子:
--
通讯作者:
Kierans, Aidan
Kierans, Aidan
中科院分区:
--
文献类型:
--
作者:
Kierans, Aidan

文献摘要

参考文献

相似文献

当今的人工智能(AI)系统严重依赖人工神经网络(ANN),但其黑箱性质导致了灾难性故障和伤害的风险。为了促进可验证安全的人工智能,我的研究将从博弈论的角度确定激励的约束,将这些约束与知识图所表示的道德知识联系起来,并揭示神经模型如何通过新的可解释性方法满足这些约束。具体地说,我将开发通过预测和隔离模型的目标来描述模型的决策过程的技术,特别是与来自知识图的值相关的目标。我的研究将允许对关键的人工智能系统进行审计,以服务于有效的监管。
Today’s artificial intelligence (AI) systems rely heavily on Artificial Neural Networks (ANNs), yet their black box nature induces risk of catastrophic failure and harm. In order to promote verifiably safe AI, my research will determine constraints on incentives from a game-theoretic perspective, tie those constraints to moral knowledge as represented by a knowledge graph, and reveal how neural models meet those constraints with novel interpretability methods. Specifically, I will develop techniques for describing models’ decision-making processes by predicting and isolating their goals, especially in relation to values derived from knowledge graphs. My research will allow critical AI systems to be audited in service of effective regulation.
DOI: --
发表时间: 2021
影响因子: 3.5
作者:
Liwei Jiang;Chandra Bhagavatula;Jenny T Liang;Jesse Dodge;Keisuke Sakaguchi;Maxwell Forbes;Jon Borchardt;Saadia Gabriel;Yulia Tsvetkov;Regina A. Rini;Yejin Choi
通讯作者: Liwei Jiang;Chandra Bhagavatula;Jenny T Liang;Jesse Dodge;Keisuke Sakaguchi;Maxwell Forbes;Jon Borchardt;Saadia Gabriel;Yulia Tsvetkov;Regina A. Rini;Yejin Choi