Benchmarked Ethics: A Roadmap to AI Alignment, Moral Knowledge, and Control
Benchmarked Ethics: A Roadmap to AI Alignment, Moral Knowledge, and Control
复制标题
道德基准:人工智能联盟、道德知识和控制的路线图
DOI:
10.1145/3600211.3604764
复制
发表时间:
2023
期刊:
影响因子:
--
通讯作者:
Kierans, Aidan
中科院分区:
文献类型:
--
作者:
Kierans, Aidan
Today’s artificial intelligence (AI) systems rely heavily on Artificial Neural Networks (ANNs), yet their black box nature induces risk of catastrophic failure and harm. In order to promote verifiably safe AI, my research will determine constraints on incentives from a game-theoretic perspective, tie those constraints to moral knowledge as represented by a knowledge graph, and reveal how neural models meet those constraints with novel interpretability methods. Specifically, I will develop techniques for describing models’ decision-making processes by predicting and isolating their goals, especially in relation to values derived from knowledge graphs. My research will allow critical AI systems to be audited in service of effective regulation.
影响因子:
3.5
作者:
Liwei Jiang;Chandra Bhagavatula;Jenny T Liang;Jesse Dodge;Keisuke Sakaguchi;Maxwell Forbes;Jon Borchardt;Saadia Gabriel;Yulia Tsvetkov;Regina A. Rini;Yejin Choi
通讯作者:
Liwei Jiang;Chandra Bhagavatula;Jenny T Liang;Jesse Dodge;Keisuke Sakaguchi;Maxwell Forbes;Jon Borchardt;Saadia Gabriel;Yulia Tsvetkov;Regina A. Rini;Yejin Choi