Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units

Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units
复制标题

DOI:
10.1186/s12911-020-1120-5
复制
发表时间:
2020-07-09
影响因子:
3.5
通讯作者:
Dong, Yinzhao
Dong, Yinzhao
中科院分区:
医学3区
文献类型:
--
作者:
Yu, Chao;Ren, Guoqi;Dong, Yinzhao

文献摘要

被引文献

相似文献

背景强化学习(RL)提供了一种有前景的技术来解决医疗保健领域复杂的顺序决策问题。近年来,应用强化学习在解决重症监护病房 (ICU) 决策问题方面取得了巨大进展。然而,由于传统强化学习算法的目标是最大化长期奖励函数,学习过程中的探索可能会对患者产生致命的影响。因此,还应该考虑短期目标,以在治疗过程中保持患者的稳定。方法我们使用监督-演员-批评家(SAC)强化学习算法,将强化学习的长期目标导向特性与监督学习的短期目标相结合来解决这个问题。我们评估了SAC与传统Actor-Critic(AC)算法在解决ICU通气和镇静剂量决策问题上的差异。结果结果表明,SAC算法在收敛速度和数据利用率方面比传统AC算法要高效得多。结论SAC算法不仅旨在长期治愈患者,而且减少了与临床医生所采用策略的偏差程度,从而提高了治疗效果。
BackgroundReinforcement learning (RL) provides a promising technique to solve complex sequential decision making problems in healthcare domains. Recent years have seen a great progress of applying RL in addressing decision-making problems in Intensive Care Units (ICUs). However, since the goal of traditional RL algorithms is to maximize a long-term reward function, exploration in the learning process may have a fatal impact on the patient. As such, a short-term goal should also be considered to keep the patient stable during the treating process.MethodsWe use a Supervised-Actor-Critic (SAC) RL algorithm to address this problem by combining the long-term goal-oriented characteristics of RL with the short-term goal of supervised learning. We evaluate the differences between SAC and traditional Actor-Critic (AC) algorithms in addressing the decision making problems of ventilation and sedative dosing in ICUs.ResultsResults show that SAC is much more efficient than the traditional AC algorithm in terms of convergence rate and data utilization.ConclusionsThe SAC algorithm not only aims to cure patients in the long term, but also reduces the degree of deviation from the strategy applied by clinical doctors and thus improves the therapeutic effect.