Diagnosis of pharyngeal cancer on endoscopic video images by Mask region-based convolutional neural network

Diagnosis of pharyngeal cancer on endoscopic video images by Mask region-based convolutional neural network
复制标题

DOI:
10.1111/den.13800
复制
发表时间:
2020-09-16
影响因子:
5.3
通讯作者:
Tada, Tomohiro
Tada, Tomohiro
中科院分区:
医学2区
文献类型:
--
作者:
Kono, Mitsuhiro;Ishihara, Ryu;Tada, Tomohiro

文献摘要

被引文献

相似文献

目的开发一种人工智能(AI)系统,用于实时诊断咽癌。方法收集在我院治疗的咽癌患者的内镜视频图像和静态图像。来自276名患者的总共4559张经病理证实的咽癌图像(1243张使用白色光成像,3316张使用窄带成像/蓝色激光成像)被用作训练数据集。人工智能系统使用了卷积神经网络(CNN)模型,这是用于分析视觉图像的典型模型。监督学习用于训练CNN。使用在我们医院拍摄的25张咽癌视频图像和36张正常咽部视频图像的独立验证数据集对AI系统进行了评估。结果人工智能系统诊断咽癌23/25例(92%),非癌17/36例(47%)。人工智能系统的处理速度为0.03秒/张图像,满足实时诊断所需的速度。对癌症检测的敏感性、特异性和准确性分别为92%、47%和66%。结论我们的单机构研究表明,我们的人工智能系统诊断咽部癌症具有良好的性能,具有高灵敏度和可接受的特异性。需要对包括多个中心的更大数据集进行系统的进一步训练和改进。
Objectives We aimed to develop an artificial intelligence (AI) system for the real-time diagnosis of pharyngeal cancers. Methods Endoscopic video images and still images of pharyngeal cancer treated in our facility were collected. A total of 4559 images of pathologically proven pharyngeal cancer (1243 using white light imaging and 3316 using narrow-band imaging/blue laser imaging) from 276 patients were used as a training dataset. The AI system used a convolutional neural network (CNN) model typical of the type used to analyze visual imagery. Supervised learning was used to train the CNN. The AI system was evaluated using an independent validation dataset of 25 video images of pharyngeal cancer and 36 video images of normal pharynx taken at our hospital. Results The AI system diagnosed 23/25 (92%) pharyngeal cancers as cancers and 17/36 (47%) non-cancers as non-cancers. The transaction speed of the AI system was 0.03 s per image, which meets the required speed for real-time diagnosis. The sensitivity, specificity, and accuracy for the detection of cancer were 92%, 47%, and 66% respectively. Conclusions Our single-institution study showed that our AI system for diagnosing cancers of the pharyngeal region had promising performance with high sensitivity and acceptable specificity. Further training and improvement of the system are required with a larger dataset including multiple centers.