Effect of Dataset Size and Medical Image Modality on Convolutional Neural Network Model Performance for Automated Segmentation: A CT and MR Renal Tumor Imaging Study.

Effect of Dataset Size and Medical Image Modality on Convolutional Neural Network Model Performance for Automated Segmentation: A CT and MR Renal Tumor Imaging Study.
复制标题

DOI:
10.1007/s10278-023-00804-1
复制
发表时间:
2023-08
影响因子:
4.4
通讯作者:
Kline, Timothy L.
Kline, Timothy L.
中科院分区:
工程技术2区
文献类型:
--
作者:
Gottlich, Harrison C.;Gregory, Adriana V.;Sharma, Vidit;Khanna, Abhinav;Moustafa, Amr U.;Lohse, Christine M.;Potretzke, Theodora A.;Korfiatis, Panagiotis;Potretzke, Aaron M.;Denic, Aleksandar;Rule, Andrew D.;Takahashi, Naoki;Erickson, Bradley J.;Leibovich, Bradley C.;Kline, Timothy L.

文献摘要

参考文献

被引文献

相似文献

本研究的目的是研究使用指数平台模型来确定产生最大医学图像分割性能所需的训练数据集大小。从我们的肾切除登记处回顾性收集 1997 年至 2017 年间获得的肾肿瘤患者的 CT 和 MR 图像。将包含 50、100、150、200、250 和 300 张图像的基于模态的数据集组装起来,以训练模型,并根据 50 个随机保留的测试集图像进行评估,并进行 80-20 个训练验证分割。使用 KiTS21 数据集的第三个实验也用于探索不同模型架构的效果。指数平台模型用于建立数据集大小与模型泛化性能的关系。为了在 CT 和 MR 成像上分割非肿瘤性肾脏区域,我们的模型产生了测试 Dice 评分平台,达到平台所需的训练验证图像数量分别为 54 和 122。为了分割 CT 和 MR 肿瘤区域,我们对 和 的测试 Dice 评分平台进行了建模,需要 125 和 389 个训练验证图像才能达到平台。对于 KiTS21 数据集,nn-UNet 2D 和 3D 架构的最佳 Dice 得分平台为 177 和 440,达到性能平台的数量为 177 和 440。我们的研究验证了不同的成像模式、目标结构和模型架构都会影响达到性能平台所需的训练图像数量。我们开发的建模方法将帮助未来的研究人员确定他们的实验何时额外的训练验证图像可能不会进一步提高模型性能。
The aim of this study is to investigate the use of an exponential-plateau model to determine the required training dataset size that yields the maximum medical image segmentation performance. CT and MR images of patients with renal tumors acquired between 1997 and 2017 were retrospectively collected from our nephrectomy registry. Modality-based datasets of 50, 100, 150, 200, 250, and 300 images were assembled to train models with an 80–20 training-validation split evaluated against 50 randomly held out test set images. A third experiment using the KiTS21 dataset was also used to explore the effects of different model architectures. Exponential-plateau models were used to establish the relationship of dataset size to model generalizability performance. For segmenting non-neoplastic kidney regions on CT and MR imaging, our model yielded test Dice score plateaus of and with the number of training-validation images needed to reach the plateaus of 54 and 122, respectively. For segmenting CT and MR tumor regions, we modeled a test Dice score plateau of and , with 125 and 389 training-validation images needed to reach the plateaus. For the KiTS21 dataset, the best Dice score plateaus for nn-UNet 2D and 3D architectures were and with number to reach performance plateau of 177 and 440. Our research validates that differing imaging modalities, target structures, and model architectures all affect the amount of training images required to reach a performance plateau. The modeling approach we developed will help future researchers determine for their experiments when additional training-validation images will likely not further improve model performance.
DOI: 10.1016/j.phro.2022.01.002
发表时间: 2022-01
影响因子: --
作者:
Lappas G;Staut N;Lieuwes NG;Biemans R;Wolfs CJA;van Hoof SJ;Dubois LJ;Verhaegen F
通讯作者: Verhaegen F
DOI: 10.1681/asn.2020040449
发表时间: 2020-11-01
影响因子: 13.6
作者:
Denic, Aleksandar;Elsherbiny, Hisham;Rule, Andrew D.
通讯作者: Rule, Andrew D.
DOI: 10.1109/tpami.2021.3100536
发表时间: 2022-10-01
影响因子: 23.6
作者:
Ma, Jun;Zhang, Yao;Yang, Xiaoping
通讯作者: Yang, Xiaoping
DOI: 10.1007/s00330-018-5695-5
发表时间: 2019-03-01
期刊: EUROPEAN RADIOLOGY
影响因子: 5.9
作者:
Joskowicz, Leo;Cohen, D.;Sosna, J.
通讯作者: Sosna, J.
DOI: 10.1117/1.jmi.6.3.034001
发表时间: 2019-07-01
影响因子: 2.4
作者:
Mueller, Sabine;Farag, Iva;Graf, Norbert
通讯作者: Graf, Norbert