Understanding machine learning classifier decisions in automated radiotherapy quality assurance

Understanding machine learning classifier decisions in automated radiotherapy quality assurance
复制标题

DOI:
10.1088/1361-6560/ac3e0e
复制
发表时间:
2022-01-21
影响因子:
3.5
通讯作者:
McIntosh, Chris
McIntosh, Chris
中科院分区:
工程技术2区
文献类型:
--
作者:
Chen, Yunsheng;Aleman, Dionne M.;McIntosh, Chris

文献摘要

被引文献

相似文献

产生放射治疗的复杂性要求严格的质量保证(QA)过程,以确保患者安全并避免临床重大错误。机器学习分类器已被探索,以扩大传统放疗治疗计划QA过程的范围和效率。然而,依赖分类器进行放疗治疗计划QA的一个重要缺陷是缺乏对具体分类器预测背后的理解。我们开发了解释方法来理解两个自动QA分类器的决策:(1)感兴趣区域(ROI)分割/标记分类器,和(2)治疗计划接受分类器。对于每个分类器,构建了一个局部可解释模型不可知解释(LIME)框架和一个新的基于团队的Shapley值框架。我们在两个放疗治疗部位(前列腺和乳房)的数据集中测试了这些方法,并证明了使用可解释的机器学习方法评估QA分类器的重要性。我们还提出了解释一致性的概念来评估分类器的性能。我们的解释方法可以很容易地可视化和人类专家评估放射治疗QA中的分类器决策。值得注意的是,我们发现基于团队的Shapley方法比LIME更具一致性。解释和验证自动决策的能力在医学治疗中至关重要。这个分析让我们得出结论,这两个QA分类器都是适度可信的,可以用来确认专家的决定,尽管当前的QA分类器不应该被视为人类QA过程的替代品。
The complexity of generating radiotherapy treatments demands a rigorous quality assurance (QA) process to ensure patient safety and to avoid clinically significant errors. Machine learning classifiers have been explored to augment the scope and efficiency of the traditional radiotherapy treatment planning QA process. However, one important gap in relying on classifiers for QA of radiotherapy treatment plans is the lack of understanding behind a specific classifier prediction. We develop explanation methods to understand the decisions of two automated QA classifiers: (1) a region of interest (ROI) segmentation/labeling classifier, and (2) a treatment plan acceptance classifier. For each classifier, a local interpretable model-agnostic explanation (LIME) framework and a novel adaption of team-based Shapley values framework are constructed. We test these methods in datasets for two radiotherapy treatment sites (prostate and breast), and demonstrate the importance of evaluating QA classifiers using interpretable machine learning approaches. We additionally develop a notion of explanation consistency to assess classifier performance. Our explanation method allows for easy visualization and human expert assessment of classifier decisions in radiotherapy QA. Notably, we find that our team-based Shapley approach is more consistent than LIME. The ability to explain and validate automated decision-making is critical in medical treatments. This analysis allows us to conclude that both QA classifiers are moderately trustworthy and can be used to confirm expert decisions, though the current QA classifiers should not be viewed as a replacement for the human QA process.