Model, Data and Reward Repair: Trusted Machine Learning for Markov Decision Processes

Model, Data and Reward Repair: Trusted Machine Learning for Markov Decision Processes
复制标题

DOI:
10.1109/dsn-w.2018.00064
复制
发表时间:
2018-06
期刊:
2018 48th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W)
影响因子:
--
通讯作者:
Shalini Ghosh-;Susmit Jha;A. Tiwari;P. Lincoln;Xiaojin Zhu
Shalini Ghosh-;Susmit Jha;A. Tiwari;P. Lincoln;Xiaojin Zhu
中科院分区:
其他
文献类型:
--
作者:
Shalini Ghosh-;Susmit Jha;A. Tiwari;P. Lincoln;Xiaojin Zhu

文献摘要

被引文献

相似文献

当机器学习(ML)模型用于安全至关重要或关键任务应用(例如,自动驾驶汽车,网络安全性,外科机器人技术)时,重要的是要确保它们提供一些高级保证(例如,安全性, livesice)。我们引入了一个名为“可信机学习”范式,以使ML模型更具值得信赖。我们将马尔可夫决策过程(MDP)用作基础动力学模型,并概述了三种TML方法:(1)模型修复,其中我们直接修改了学习模型; (2)数据修复,其中我们修改数据,以便从修改后的数据重新学习导致可信赖的模型; (3)奖励维修,其中我们修改MDP的奖励功能以满足指定的逻辑约束。我们展示了如何在某些适当的逻辑片段中表达所需属性(例如PCTL(例如PCTL,即概率计算树逻辑),一阶逻辑或一阶逻辑或一阶逻辑或一阶逻辑或一阶逻辑,如何对概率模型(例如MDP)进行有效进行这些维修。命题逻辑。我们说明了来自多个域的案例研究的方法,例如,避免障碍物的汽车控制器以及无线传感器网络中的查询路由控制器。
When machine learning (ML) models are used in safety-critical or mission-critical applications (e.g., self driving cars, cyber security, surgical robotics), it is important to ensure that they provide some high-level guarantees (e.g., safety, liveness). We introduce a paradigm called Trusted Machine Learning (TML) for making ML models more trustworthy. We use Markov Decision Processes (MDPs) as the underlying dynamical model and outline three TML approaches: (1) Model Repair, wherein we modify the learned model directly; (2) Data Repair, wherein we modify the data so that re-learning from the modified data results in a trusted model; and (3) Reward Repair, wherein we modify the reward function of the MDP to satisfy the specified logical constraint. We show how these repairs can be done efficiently for probabilistic models (e.g., MDP) when the desired properties are expressed in some appropriate fragment of logic such as temporal logic (for example PCTL, i.e., Probabilistic Computation Tree Logic), first order logic or propositional logic. We illustrate our approaches on case studies from multiple domains, e.g., car controller for obstacle avoidance, and a query routing controller in a wireless sensor network.