Contouring quality assurance methodology based on multiple geometric features against deep learning auto-segmentation.

Contouring quality assurance methodology based on multiple geometric features against deep learning auto-segmentation.
复制标题

基于多个几何特征和深度学习自动分割的轮廓质量保证方法。

DOI:
10.1002/mp.16299
复制
发表时间:
2023
期刊:
影响因子:
3.8
通讯作者:
Chen,Quan
Chen,Quan
中科院分区:
医学3区
文献类型:
--
作者:
Duan,Jingwei;Bernard,MarkE;Castle,JamesR;Feng,Xue;Wang,Chi;Kenamond,MarkC;Chen,Quan

文献摘要

被引文献

相似文献

背景轮廓误差是放射治疗中常见的失效模式之一。已经做出了多种努力来开发自动检测分割错误的工具。基于深度学习的自动分割(DLAS)已被用作标记人工分割错误的基线,但这些努力仅限于使用一到两个轮廓比较度量.本研究的目的是开发一种改进的轮廓质量保证系统来识别和标记手动轮廓错误.方法和材料以DLAS轮廓为参照,与人工分割的轮廓进行比较.通过对两种分割方法的比较,共确定了27个几何一致性度量。进行特征选择以优化机器学习分类模型的训练,以识别潜在的轮廓误差。使用包含339个案例的公共数据集来训练和测试该分类器。四个独立的分类器使用五次交叉验证进行训练,每个分类器的预测使用软投票进行集成。训练好的模型在一个坚持测试的数据集上得到了验证。使用另外一个包含60个病例的独立临床数据集来检验该模型的泛化能力。结果提出的机器学习多特征(ML-MF)方法优于传统的基于非机器学习的方法,这些方法只基于一个或两个几何一致性度量。对于脑干、腮腺_L、腮腺_R和下颌骨轮廓,机器学习模型的召回率分别为0.842(0.899)、0.762(0.762)、0.727(0.842)和0.773(0.773),而仅基于骰子相似系数的方法的召回率(准确率)为0.526(0.909)、0.619(0.765)、0.682(0.882)、0.773(0.568)。在外部验证数据集中,脑干、腮腺_L、腮腺_R和下颌骨轮廓的专家分别确认了66.7、93.3、94.1和58.8%的标记病例存在轮廓错误。结论与传统方法相比,提出的ML-MF方法具有更好的性能,该方法包含多个几何一致度量来标记手动轮廓错误。这种方法易于在临床实践中实施,并有助于减少与手动分割和审查相关的大量时间和人力成本。
BackgroundContouring error is one of the top failure modes in radiation treatment. Multiple efforts have been made to develop tools to automatically detect segmentation errors. Deep learning‐based auto‐segmentation (DLAS) has been used as a baseline for flagging manual segmentation errors, but those efforts are limited to using only one or two contour comparison metrics.PurposeThe purpose of this research is to develop an improved contouring quality assurance system to identify and flag manual contouring errors.Methods and materialsDLAS contours were used as a reference to compare with manually segmented contours. A total of 27 geometric agreement metrics were determined from the comparisons between the two segmentation approaches. Feature selection was performed to optimize the training of a machine learning classification model to identify potential contouring errors. A public dataset with 339 cases was used to train and test the classifier. Four independent classifiers were trained using five‐fold cross validation, and the predictions from each classifier were ensembled using soft voting. The trained model was validated on a held‐out testing dataset. An additional independent clinical dataset with 60 cases was used to test the generalizability of the model. Model predictions were reviewed by an expert to confirm or reject the findings.ResultsThe proposed machine learning multiple features (ML‐MF) approach outperformed traditional nonmachine‐learning‐based approaches that are based on only one or two geometric agreement metrics. The machine learning model achieved recall (precision) values of 0.842 (0.899), 0.762 (0.762), 0.727 (0.842), and 0.773 (0.773) for Brainstem, Parotid_L, Parotid_R, and mandible contours, respectively compared to 0.526 (0.909), 0.619 (0.765), 0.682 (0.882), 0.773 (0.568) for an approach based solely on Dice similarity coefficient values. In the external validation dataset, 66.7, 93.3, 94.1, and 58.8% of flagged cases were confirmed to have contouring errors by an expert for Brainstem, Parotid_L, Parotid_R, and mandible contours, respectively.ConclusionsThe proposed ML‐MF approach, which includes multiple geometric agreement metrics to flag manual contouring errors, demonstrated superior performance in comparison to traditional methods. This method is easy to implement in clinical practice and can help to reduce the significant time and labor costs associated with manual segmentation and review.