Decision making with limited feedback: Error bounds for predictive policing and recidivism prediction

Decision making with limited feedback: Error bounds for predictive policing and recidivism prediction
复制标题

DOI:
--
复制
发表时间:
2018
期刊:
Engineering and Technology Journal
影响因子:
--
通讯作者:
D. Ensign;Sorelle A. Frielder;Scott Neville;Carlos Scheidegger;Suresh Venkatasubramanian;M. Mohri;Karthik Sridharan
D. Ensign;Sorelle A. Frielder;Scott Neville;Carlos Scheidegger;Suresh Venkatasubramanian;M. Mohri;Karthik Sridharan
中科院分区:
其他
文献类型:
--
作者:
D. Ensign;Sorelle A. Frielder;Scott Neville;Carlos Scheidegger;Suresh Venkatasubramanian;M. Mohri;Karthik Sridharan

文献摘要

相似文献

当对模型进行训练以便在各种现实环境中部署决策时,它们通常以批处理模式进行训练。历史数据用于在部署之前训练和验证模型。然而,在许多情况下,反馈改变了培训过程的性质。要么是学习者没有得到关于其行为的完整反馈,要么是经过训练的模型所做的决定影响了它将看到的未来训练数据。本文主要研究累犯预测和预测性警务问题。我们通过展示这两个问题(以及其他类似的问题)可以抽象成一个称为部分监控的通用强化学习框架,提出了第一个具有可证明遗憾的算法。我们还讨论了这些解决方案的政策含义。
When models are trained for deployment in decision-making in various real-world settings, they are typically trained in batch mode. Historical data is used to train and validate the models prior to deployment. However, in many settings, feedback changes the nature of the training process. Either the learner does not get full feedback on its actions, or the decisions made by the trained model influence what future training data it will see. In this paper, we focus on the problems of recidivism prediction and predictive policing. We present the first algorithms with provable regret for these problems, by showing that both problems (and others like these) can be abstracted into a general reinforcement learning framework called partial monitoring. We also discuss the policy implications of these solutions.