Methods for Reinforcement Learning in Clinical Decision Support

Methods for Reinforcement Learning in Clinical Decision Support
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Niranjani Prasad
Niranjani Prasad
中科院分区:
其他
文献类型:
--
作者:
Niranjani Prasad

文献摘要

被引文献

相似文献

从呼吸支持到疼痛管理的常规干预管理构成了住院护理的主要部分。正确的治疗对于改善患者预后和降低成本至关重要,但这些干预措施往往知之甚少,临床对最佳方案的意见可能差异很大。通过一系列关键的危重病护理干预的案例研究,本论文开发了一个临床医生在回路决策支持的框架。其中第一个探讨了患者从机械通气中的脱机:入院被建模为马尔可夫决策过程(MDP),采用无模型批量强化学习算法来学习镇静和呼吸机支持的个性化方案,当根据当前临床实践进行评估时,这些方案有望改善结果。本论文的第二部分是针对有效的奖励设计时,制定临床决策作为一个强化学习任务。在解决重症监护中的冗余测试问题时,帕累托最优强化学习方法与已知的程序约束相结合,以巩固多个往往相互冲突的临床目标,并产生灵活的优化排序策略。这里的挑战进一步探讨,以研究如何由护理提供者的决定,如在现有的数据中观察到的,可以用来限制可能的凸组合的目标在奖励函数,那些产生的政策,反映了我们隐含地知道从数据的合理行为的任务,并允许高信心的政策评估。所提出的方法来奖励设计证明,通过合成域,以及在重症监护规划。最后一个案例研究考虑了电解质补充的任务,描述了如何使用MDP框架优化这项任务,并通过强化学习的透镜分析当前的临床行为,然后概述了在当前医疗保健系统中采用这些工具所需的步骤。
The administration of routine interventions, from breathing support to pain management, constitutes a major part of inpatient care. Thoughtful treatment is crucial to improving patient outcomes and minimizing costs, but these interventions are often poorly understood, and clinical opinion on best protocols can vary significantly. Through a series of case studies of key critical care interventions, this thesis develops a framework for clinician-in-loop decision support. The first of these explores the weaning of patients from mechanical ventilation: admissions are modelled as Markov decision processes (MDPs), and model-free batch reinforcement learning algorithms are employed to learn personalized regimes of sedation and ventilator support, that show promise in improving outcomes when assessed against current clinical practice. The second part of this thesis is directed towards effective reward design when formulating clinical decisions as a reinforcement learning task. In tackling the problem of redundant testing in critical care, methods for Pareto-optimal reinforcement learning are integrated with known procedural constraints in order to consolidate multiple, often conflicting, clinical goals and produce a flexible optimized ordering policy. The challenges here are probed further to examine how decisions by care providers, as observed in available data, can be used to restrict the possible convex combinations of objectives in the reward function, to those that yield policies reflecting what we implicitly know from the data about reasonable behaviour for a task, and that allow for high-confidence off-policy evaluation. The proposed approach to reward design is demonstrated through synthetic domains as well as in planning in critical care. The final case study considers the task of electrolyte repletion, describing how this task can be optimized using the MDP framework and analysing current clinical behaviour through the lens of reinforcement learning, before going on to outline the steps necessary in enabling the adoption of these tools in current healthcare systems.