Counterfactual harm

Counterfactual harm
复制标题

DOI:
--
复制
发表时间:
2022-04
期刊:
--
影响因子:
--
通讯作者:
Jonathan G. Richens;R. Beard;Daniel H. Thompson
Jonathan G. Richens;R. Beard;Daniel H. Thompson
中科院分区:
其他
文献类型:
--
作者:
Jonathan G. Richens;R. Beard;Daniel H. Thompson

文献摘要

被引文献

相似文献

为了在现实世界中安全和合乎道德地行事,代理人必须能够对伤害进行推理,并避免有害行为。然而,到目前为止,还没有统计方法来衡量危害并将其纳入算法决策。在这篇文章中,我们提出了第一个使用因果模型的危害和收益的正式定义。我们证明了在某些情况下,任何对伤害的事实定义都必须违反基本的直觉,并证明了不能执行反事实推理的标准机器学习算法保证在分布转移之后追求有害策略。我们使用我们对伤害的定义来设计一个框架,用于使用反事实的目标函数进行厌恶伤害的决策。我们使用从随机对照试验数据中学习的剂量-反应模型,在确定最佳药物剂量的问题上展示了该框架。我们发现,根据治疗效果选择剂量的标准方法会导致不必要的有害剂量,而我们的反事实方法允许我们在不牺牲疗效的情况下确定危害显著较小的剂量。
To act safely and ethically in the real world, agents must be able to reason about harm and avoid harmful actions. However, to date there is no statistical method for measuring harm and factoring it into algorithmic decisions. In this paper we propose the first formal definition of harm and benefit using causal models. We show that any factual definition of harm must violate basic intuitions in certain scenarios, and show that standard machine learning algorithms that cannot perform counterfactual reasoning are guaranteed to pursue harmful policies following distributional shifts. We use our definition of harm to devise a framework for harm-averse decision making using counterfactual objective functions. We demonstrate this framework on the problem of identifying optimal drug doses using a dose-response model learned from randomized control trial data. We find that the standard method of selecting doses using treatment effects results in unnecessarily harmful doses, while our counterfactual approach allows us to identify doses that are significantly less harmful without sacrificing efficacy.