Automatic classification of ultrasound breast lesions using a deep convolutional neural network mimicking human decision-making

Automatic classification of ultrasound breast lesions using a deep convolutional neural network mimicking human decision-making
复制标题

DOI:
10.1007/s00330-019-06118-7
复制
发表时间:
2019-10-01
期刊:
影响因子:
5.9
通讯作者:
Boss, Andreas
Boss, Andreas
中科院分区:
医学2区
文献类型:
--
作者:
Ciritsis, Alexander;Rossi, Cristina;Boss, Andreas

文献摘要

被引文献

相似文献

目的根据乳腺成像报告和数据系统(BI-RADS),评估用于检测、突出显示和分类超声(US)乳腺病变的深度卷积神经网络(dCNN),以模仿人类决策。方法和材料从582例患者(年龄56.3 +/- 11.5岁)的1019乳腺超声图像链接到相应的放射学报告。病变分为以下类别:无组织、正常乳腺组织、BI-RADS 2(囊肿、淋巴结)、BI-RADS 3(非囊性肿块)和BI-RADS 4-5(可疑)。为了测试dCNN的准确性,由dCNN和两名独立的读取器评估了一个内部数据集(101张图像)和一个外部测试数据集(43张图像)。放射学报告、组织病理学结果和随访检查作为参考。dCNN和人类的性能在分类准确性和受试者工作特征(ROC)曲线方面进行了量化。结果在内部测试数据集中,dCNN区分BI-RADS 2和BI-RADS 3-5病变的分类准确率为87.1%(外部93.0%),而人类读者的分类准确率为79.2 +/- 1.9%(外部95.3 +/- 2.3%)。对于BI-RADS 2-3与BI-RADS 4-5的分类,dCNN达到了93.1%的分类准确度(外部95.3%),而人类的分类准确度为91.6 +/- 5.4%(外部94.1 +/- 1.2%)。dCNN的内部数据集的AUC为83.8(外部96.7),人类的AUC为84.6 +/- 2.3(外部90.9 +/- 2.9)。结论dCNN可用于模拟人类决策,根据BI-RADS目录评估乳腺病变的单个US图像。该技术达到高精度,并可用于标准化的高度依赖于超声评估。
Objectives To evaluate a deep convolutional neural network (dCNN) for detection, highlighting, and classification of ultrasound (US) breast lesions mimicking human decision-making according to the Breast Imaging Reporting and Data System (BI-RADS). Methods and materials One thousand nineteen breast ultrasound images from 582 patients (age 56.3 +/- 11.5 years) were linked to the corresponding radiological report. Lesions were categorized into the following classes: no tissue, normal breast tissue, BI-RADS 2 (cysts, lymph nodes), BI-RADS 3 (non-cystic mass), and BI-RADS 4-5 (suspicious). To test the accuracy of the dCNN, one internal dataset (101 images) and one external test dataset (43 images) were evaluated by the dCNN and two independent readers. Radiological reports, histopathological results, and follow-up examinations served as reference. The performances of the dCNN and the humans were quantified in terms of classification accuracies and receiver operating characteristic (ROC) curves. Results In the internal test dataset, the classification accuracy of the dCNN differentiating BI-RADS 2 from BI-RADS 3-5 lesions was 87.1% (external 93.0%) compared with that of human readers with 79.2 +/- 1.9% (external 95.3 +/- 2.3%). For the classification of BI-RADS 2-3 versus BI-RADS 4-5, the dCNN reached a classification accuracy of 93.1% (external 95.3%), whereas the classification accuracy of humans yielded 91.6 +/- 5.4% (external 94.1 +/- 1.2%). The AUC on the internal dataset was 83.8 (external 96.7) for the dCNN and 84.6 +/- 2.3 (external 90.9 +/- 2.9) for the humans. Conclusion dCNNs may be used to mimic human decision-making in the evaluation of single US images of breast lesion according to the BI-RADS catalog. The technique reaches high accuracies and may serve for standardization of highly observer-dependent US assessment.