Deconfounding Actor-Critic Network with Policy Adaptation for Dynamic Treatment Regimes.

Deconfounding Actor-Critic Network with Policy Adaptation for Dynamic Treatment Regimes.
复制标题

DOI:
10.1145/3534678.3539413
复制
发表时间:
2022-08
期刊:
KDD : proceedings. International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

尽管在基础和临床研究方面做出了巨大努力,但针对危重患者的个性化通气策略仍然是一个重大挑战。最近,在电子健康记录(EHR)上使用强化学习(RL)的动态治疗方案(DTR)引起了医疗保健行业和机器学习研究界的兴趣。然而,由于混杂因素的存在,大多数已知的DTR政策可能会有偏见。虽然非幸存者接受的一些治疗措施可能有帮助,但如果混杂因素导致死亡,则由长期结果指导的RL模型的训练(例如,90-日死亡率)将惩罚那些导致所学习的DTR策略次优的治疗动作。在这项研究中,我们开发了一个新的去基础演员-评论家网络(DAC),以学习最佳的DTR政策的患者。为了减轻混淆的问题,我们将病人的呼吸模块和混淆的平衡模块到我们的演员批评框架。为了避免惩罚非幸存者接受的有效治疗行动,我们设计了一个短期奖励,以捕捉患者的即时健康状态变化。将短期和长期奖励相结合可以进一步提高模型的性能。此外,我们引入了一种策略自适应方法,成功地将学习模型转移到新的小规模数据集。在一个半合成数据集和两个不同的真实数据集上的实验结果表明,该模型的性能优于现有的模型。所提出的模型为机械通气提供了个性化的治疗决策,可以改善患者的预后。
Despite intense efforts in basic and clinical research, an individualized ventilation strategy for critically ill patients remains a major challenge. Recently, dynamic treatment regime (DTR) with reinforcement learning (RL) on electronic health records (EHR) has attracted interest from both the healthcare industry and machine learning research community. However, most learned DTR policies might be biased due to the existence of confounders. Although some treatment actions non-survivors received may be helpful, if confounders cause the mortality, the training of RL models guided by long-term outcomes (e.g., 90-day mortality) would punish those treatment actions causing the learned DTR policies to be suboptimal. In this study, we develop a new deconfounding actor-critic network (DAC) to learn optimal DTR policies for patients. To alleviate confounding issues, we incorporate a patient resampling module and a confounding balance module into our actor-critic framework. To avoid punishing the effective treatment actions non-survivors received, we design a short-term reward to capture patients’ immediate health state changes. Combining short-term with long-term rewards could further improve the model performance. Moreover, we introduce a policy adaptation method to successfully transfer the learned model to new-source small-scale datasets. The experimental results on one semi-synthetic and two different real-world datasets show the proposed model outperforms the state-of-the-art models. The proposed model provides individualized treatment decisions for mechanical ventilation that could improve patient outcomes.