Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer.

Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer.
复制标题

DOI:
10.1111/j.1541-0420.2011.01572.x
复制
发表时间:
2011-12
期刊:
影响因子:
1.9
通讯作者:
Kosorok MR
Kosorok MR
中科院分区:
数学3区
文献类型:
--
作者:
Zhao Y;Zeng D;Socinski MA;Kosorok MR

文献摘要

参考文献

被引文献

相似文献

晚期转移性IIIB/IV期非小细胞肺癌(NSCLC)的典型治疗方案包括多条治疗路线。我们提出了一种自适应强化学习方法,从一项专门设计的针对未接受过系统治疗的晚期非小细胞肺癌患者的实验性治疗的临床试验(临床强化试验)中发现最佳的个体化治疗方案。除了根据预后因素为一线和二线治疗选择最佳化合物的问题的复杂性外,另一个主要目标是确定启动二线治疗的最佳时间,无论是立即还是在诱导治疗后推迟,从而产生最长的总体生存时间。使用了一种称为Q-学习的强化学习方法,该方法涉及从从临床强化试验产生的患者数据中学习最佳方案。利用删失数据对支持向量机回归进行了改进,实现了用时间指标参数逼近Q-函数。在此框架内,模拟研究表明,该方法可以直接从临床数据中提取两个治疗路线的最优方案,而不需要事先知道治疗的作用机制。此外,我们证明,该设计可靠地选择了二线治疗的最佳初始时间,同时考虑了非小细胞肺癌患者之间的异质性。
Typical regimens for advanced metastatic stage IIIB/IV non-small cell lung cancer (NSCLC) consist of multiple lines of treatment. We present an adaptive reinforcement learning approach to discover optimal individualized treatment regimens from a specially designed clinical trial (a “clinical reinforcement trial”) of an experimental treatment for patients with advanced NSCLC who have not been treated previously with systemic therapy. In addition to the complexity of the problem of selecting optimal compounds for first and second-line treatments based on prognostic factors, another primary goal is to determine the optimal time to initiate second-line therapy, either immediately or delayed after induction therapy, yielding the longest overall survival time. A reinforcement learning method called Q-learning is utilized which involves learning an optimal regimen from patient data generated from the clinical reinforcement trial. Approximating the Q-function with time-indexed parameters can be achieved by using a modification of support vector regression which can utilize censored data. Within this framework, a simulation study shows that the procedure can extract optimal regimens for two lines of treatment directly from clinical data without prior knowledge of the treatment effect mechanism. In addition, we demonstrate that the design reliably selects the best initial time for second-line therapy while taking into account the heterogeneity of NSCLC across patients.
DOI: 10.1634/theoncologist.13-s1-28
发表时间: 2008-01-01
期刊: ONCOLOGIST
影响因子: 5.8
作者:
Stinchcombe, Thomas E.;Socinski, Mark A.
通讯作者: Socinski, Mark A.
DOI: 10.1056/nejmoa061884
发表时间: 2006-12-14
影响因子: 158.5
作者:
Sandler, Alan;Gray, Robert;Johnson, David H.
通讯作者: Johnson, David H.
DOI: 10.2307/2684631
发表时间: 1995-05-01
影响因子: 1.8
作者:
GELBER, RD;COLE, BF;GOLDHIRSCH, A
通讯作者: GOLDHIRSCH, A
DOI: 10.1002/sim.2894
发表时间: 2007-11-20
影响因子: 2
作者:
Thall, Peter F.;Wooten, Leiko H.;Tannir, Nizar M.
通讯作者: Tannir, Nizar M.
DOI: 10.1016/s0140-6736(09)61497-5
发表时间: 2009-10-24
期刊: LANCET
影响因子: 168.9
作者:
Ciuleanu, Tudor;Brodowicz, Thomas;Belani, Chandra P.
通讯作者: Belani, Chandra P.