Deep reinforcement learning for automated radiation adaptation in lung cancer.

Deep reinforcement learning for automated radiation adaptation in lung cancer.
复制标题

DOI:
10.1002/mp.12625
复制
发表时间:
2017-12
期刊:
影响因子:
3.8
通讯作者:
Naqa IE
Naqa IE
中科院分区:
医学3区
文献类型:
--
作者:
Tseng HH;Luo Y;Cui S;Chien JT;Ten Haken RK;Naqa IE

文献摘要

参考文献

被引文献

相似文献

研究基于历史治疗计划的深度强化学习(DRL),为非小细胞肺癌(NSCLC)患者开发自动化放射适应方案,旨在最大限度地提高肿瘤局部控制,降低2级放射性肺炎(RP 2)的发生率。在114例接受放射治疗的NSCLC患者的回顾性人群中,开发了一种3组分神经网络框架,用于剂量分割适应的深度强化学习(DRL)。大规模的患者特征包括临床,遗传和成像放射组学特征,以及肿瘤和肺剂量学变量。首先,采用生成对抗网络(GAN)从相对有限的样本量中学习DRL训练所需的患者人群特征。其次,利用原始数据和合成数据(通过GAN)通过深度神经网络(DNN)重建放射治疗人工环境(RAE),以估计用于适应个性化放射治疗患者治疗过程的转移概率。第三,将深度Q网络(DQN)应用于RAE,以在反应适应性治疗设置中选择最佳剂量。这种多组分强化学习方法是以应用于自适应剂量递增临床方案的真实的临床决策为基准的。其中,34例患者根据肿瘤中的avid PET信号进行治疗,并受到RP 2正常组织并发症概率(NTCP)限值17.2%的限制。简单治愈概率(P+)用作DRL中的基线奖励函数。以我们的自适应剂量递增方案作为所提出的DRL(GAN+RAE+DQN)架构的蓝图,我们获得了在放射治疗过程中约2/3处使用的自动剂量适应估计。通过让DQN组件自由控制估计的适应性剂量/分次(范围为1 ~ 5戈伊),DRL自动支持剂量递增/递减1.5 ~ 3.8戈伊,该范围与临床方案中使用的范围相似。相同的DQN为34名测试患者产生了两种剂量递增模式,但具有不同的奖励变体。首先,使用基线P+奖励函数,DQN的个体自适应分数剂量具有与RMSE= 0.76戈伊的临床数据相似的趋势;但DQN建议的适应通常幅度较低(侵略性较低)。其次,通过调整P+奖励函数,更强调减轻局部失效,实现了DQN和临床方案之间剂量的更好匹配,RMSE= 0.5戈伊。此外,DQN选择的决定似乎与患者的最终结果更一致。相比之下,用于强化学习的传统时间差(TD)算法由于数值不稳定性和缺乏足够的学习而产生RMSE= 3.3戈伊。我们证明了DRL的自动剂量适应是一种可行的,有前途的方法,可以实现与临床医生选择的结果相似的结果。如果要考虑个别情况,该过程可能需要定制奖励函数。然而,将该框架开发成用于临床决策支持的完全可信的自主系统将需要在更大的多机构数据集上进行进一步验证。
To investigate deep reinforcement learning (DRL) based on historical treatment plans for developing automated radiation adaptation protocols for non-small cell lung cancer (NSCLC) patients that aim to maximize tumor local control at reduced rates of radiation pneumonitis grade 2 (RP2). In a retrospective population of 114 NSCLC patients who received radiotherapy, a 3-component neural networks framework was developed for deep reinforcement learning (DRL) of dose fractionation adaptation. Large-scale patient characteristics included clinical, genetic, and imaging radiomics features in addition to tumor and lung dosimetric variables. First, a generative adversarial network (GAN) was employed to learn patient population characteristics necessary for DRL training from a relatively limited sample size. Second, a radiotherapy artificial environment (RAE) was reconstructed by a deep neural network (DNN) utilizing both original and synthetic data (by GAN) to estimate the transition probabilities for adaptation of personalized radiotherapy patients’ treatment courses. Third, a deep Q-network (DQN) was applied to the RAE for choosing the optimal dose in a response-adapted treatment setting. This multi-component reinforcement learning approach was benchmarked against real clinical decisions that were applied in an adaptive dose escalation clinical protocol. In which, 34 patients were treated based on avid PET signal in the tumor and constrained by a 17.2% normal tissue complication probability (NTCP) limit for RP2. The uncomplicated cure probability (P+) was used as a baseline reward function in the DRL. Taking our adaptive dose escalation protocol as a blueprint for the proposed DRL (GAN+RAE+DQN) architecture, we obtained an automated dose adaptation estimate for use at ~ 2/3 of the way into the radiotherapy treatment course. By letting the DQN component freely control the estimated adaptive dose per fraction (ranging from 1 ~ 5 Gy), the DRL automatically favored dose escalation/de-escalation between 1.5 ~ 3.8 Gy, a range similar to that used in the clinical protocol. The same DQN yielded two patterns of dose escalation for the 34 test patients, but with different reward variants. First, using the baseline P+ reward function, individual adaptive fraction doses of the DQN had similar tendencies to the clinical data with an RMSE= 0.76 Gy; but adaptations suggested by the DQN were generally lower in magnitude (less aggressive). Second, by adjusting the P+ reward function with higher emphasis on mitigating local failure, better matching of doses between the DQN and the clinical protocol was achieved with an RMSE= 0.5 Gy. Moreover, the decisions selected by the DQN seemed to have better concordance with patients eventual outcomes. In comparison, the traditional temporal difference (TD) algorithm for reinforcement learning yielded an RMSE= 3.3 Gy due to numerical instabilities and lack of sufficient learning. We demonstrated that automated dose adaptation by DRL is a feasible and a promising approach for achieving similar results to those chosen by clinicians. The process may require customization of the reward function if individual cases were to be considered. However, development of this framework into a fully credible autonomous system for clinical decision support would require further validation on larger multi-institutional datasets.
DOI: 10.1016/s1470-2045(14)71207-0
发表时间: 2015-02
期刊: LANCET ONCOLOGY
影响因子: 51.1
作者:
Bradley, Jeffrey D.;Paulus, Rebecca;Komaki, Ritsuko;Masters, Gregory;Blumenschein, George;Schild, Steven;Bogart, Jeffrey;Hu, Chen;Forster, Kenneth;Magliocco, Anthony;Kavadi, Vivek;Garces, Yolanda I.;Narayan, Samir;Iyengar, Puneeth;Robinson, Cliff;Wynn, Raymond B.;Koprowski, Christopher;Meng, Joanne;Beitler, Jonathan;Gaur, Rakesh;Curran, Walter, Jr.;Choy, Hak
通讯作者: Choy, Hak
DOI: 10.1016/0360-3016(90)90037-k
发表时间: 1990-10-01
影响因子: 7
作者:
AGREN, A;BRAHME, A;TURESSON, I
通讯作者: TURESSON, I
DOI: 10.1080/028418600750013113
发表时间: 2000-01-01
期刊: ACTA ONCOLOGICA
影响因子: 3.1
作者:
Bentzen, SM;Dische, S
通讯作者: Dische, S
DOI: 10.1016/0893-6080(89)90020-8
发表时间: 1989-01-01
期刊: NEURAL NETWORKS
影响因子: 7.8
作者:
HORNIK, K;STINCHCOMBE, M;WHITE, H
通讯作者: WHITE, H
DOI: 10.1088/0031-9155/54/14/007
发表时间: 2009-07-21
影响因子: 3.5
作者:
Kim, M.;Ghate, A.;Phillips, M. H.
通讯作者: Phillips, M. H.