A Markov decision process approach to temporal modulation of dose fractions in radiation therapy planning

A Markov decision process approach to temporal modulation of dose fractions in radiation therapy planning
复制标题

DOI:
10.1088/0031-9155/54/14/007
复制
发表时间:
2009-07-21
影响因子:
3.5
通讯作者:
Phillips, M. H.
Phillips, M. H.
中科院分区:
工程技术2区
文献类型:
--
作者:
Kim, M.;Ghate, A.;Phillips, M. H.

文献摘要

被引文献

相似文献

目前癌症放射治疗的最新技术在空间上优化了射束强度,使得肿瘤接受高剂量辐射,同时将对附近健康组织的损害降至最低。通常的做法是在几周内提供辐射,其中每日剂量是计划总剂量的一小部分。这种“分割方案”基于传统的放射生物学反应模型,其中正常组织细胞具有修复辐射造成的亚致死损伤的能力。这种能力在肿瘤中明显不那么突出。定量功能成像和生物标记的最新进展为测量患者在治疗过程中对辐射的反应提供了新的机会。这为设计分割时间表打开了大门,该分割时间表在确定当天的分割时考虑了患者直到特定治疗日对辐射的累积反应。我们提出了一种新颖的方法,首次从数学上探讨了此类分馏方案的好处。这是通过构建风格马尔可夫决策过程 (MDP) 模型来实现的,该模型通过状态和动作空间的直观选择以及转移概率和奖励函数合并了问题的一些关键特征。通过几个简单的数值例子探讨了该 MDP 模型的最优政策结构。
The current state of the art in cancer treatment by radiation optimizes beam intensity spatially such that tumors receive high dose radiation whereas damage to nearby healthy tissues is minimized. It is common practice to deliver the radiation over several weeks, where the daily dose is a small constant fraction of the total planned. Such a 'fractionation schedule' is based on traditional models of radiobiological response where normal tissue cells possess the ability to repair sublethal damage done by radiation. This capability is significantly less prominent in tumors. Recent advances in quantitative functional imaging and biological markers are providing new opportunities to measure patient response to radiation over the treatment course. This opens the door for designing fractionation schedules that take into account the patient's cumulative response to radiation up to a particular treatment day in determining the fraction on that day. We propose a novel approach that, for the first time, mathematically explores the benefits of such fractionation schemes. This is achieved by building a stylistic Markov decision process (MDP) model, which incorporates some key features of the problem through intuitive choices of state and action spaces, as well as transition probability and reward functions. The structure of optimal policies for this MDP model is explored through several simple numerical examples.