Cooperative Inverse Decision Theory for Uncertain Preferences

Cooperative Inverse Decision Theory for Uncertain Preferences
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Zachary Robertson;Hantao Zhang;Oluwasanmi Koyejo
Zachary Robertson;Hantao Zhang;Oluwasanmi Koyejo
中科院分区:
其他
文献类型:
--
作者:
Zachary Robertson;Hantao Zhang;Oluwasanmi Koyejo

文献摘要

相似文献

逆向决策理论(IDT)的目的是通过在实例上引出专家分类来学习分类的性能度量。然而,在实际环境中的启发可能需要许多潜在的模糊的例子的分类。为了提高启发的效率,我们提出了合作逆决策理论(CIDT)框架作为性能指标启发问题的形式化。在合作逆决策理论中,专家和机器玩一个游戏,根据专家的性能指标获得奖励,但机器最初并不知道这个函数是什么。我们表明,在这个框架中的最优策略产生主动学习,导致样本复杂性比以前的工作呈指数级改善。我们的主要发现之一是,一个广泛的类次优专家可以表示为具有不确定的偏好。我们使用这一发现来显示这些专家自然适合我们提出的框架扩展逆决策理论,以有效地处理由于噪声,冲突的专家或系统错误而导致的次优决策数据。
Inverse decision theory (IDT) aims to learn a performance metric for classification by eliciting expert classifications on examples. However, elicitation in practical settings may require many classifications of potentially ambiguous examples. To improve the efficiency of elicitation, we propose the cooperative inverse decision theory (CIDT) framework as a formalization of the performance metric elicitation problem. In cooperative inverse decision theory, the expert and a machine play a game where both are rewarded according to the expert’s performance metric, but the machine does not initially know what this function is. We show that optimal policies in this framework produce active learning that leads to an exponential improvement in sample complexity over previous work. One of our key findings is that a broad class of sub-optimal experts can be represented as having uncertain preferences. We use this finding to show such experts naturally fit into our proposed framework extending inverse decision theory to efficiently deal with decision data that is sub-optimal due to noise, conflicting experts, or systematic error.