Alignment for Advanced Machine Learning Systems

Alignment for Advanced Machine Learning Systems
复制标题

高级机器学习系统的协调

DOI:
--
复制
发表时间:
2020
期刊:
Ethics of Artificial Intelligence
影响因子:
--
通讯作者:
Andrew Critch
Andrew Critch
中科院分区:
--
文献类型:
--
作者:
Jessica Taylor;Eliezer Yudkowsky;Patrick LaVictoire;Andrew Critch

文献摘要

被引文献

相似文献

本章将围绕一个问题对八个研究领域进行调查:随着学习系统变得越来越智能化和自主化,什么样的设计原则可以最好地确保它们的行为符合操作员的利益?本章重点讨论了人工智能对齐的两个主要技术障碍:指定正确类型的目标函数的挑战,以及设计人工智能系统的挑战,即使在目标函数与设计者的意图不完全一致的情况下,也可以避免意外后果和不良行为。调查的问题包括:我们如何训练强化学习者采取行动,更适合有意义的评估,由智能监督?什么样的目标函数激励一个系统“没有太大的影响”或“没有太多的副作用”?本章讨论了这些问题、相关工作以及未来研究的潜在方向,目的是突出机器学习中今天似乎易于处理的相关研究主题。
This chapter surveys eight research areas organized around one question: As learning systems become increasingly intelligent and autonomous, what design principles can best ensure that their behavior is aligned with the interests of the operators? The chapter focuses on two major technical obstacles to AI alignment: the challenge of specifying the right kind of objective functions and the challenge of designing AI systems that avoid unintended consequences and undesirable behavior even in cases where the objective function does not line up perfectly with the intentions of the designers. The questions surveyed include the following: How can we train reinforcement learners to take actions that are more amenable to meaningful assessment by intelligent overseers? What kinds of objective functions incentivize a system to “not have an overly large impact” or “not have many side effects”? The chapter discusses these questions, related work, and potential directions for future research, with the goal of highlighting relevant research topics in machine learning that appear tractable today.