Automated grading of acne vulgaris by deep learning with convolutional neural networks

Automated grading of acne vulgaris by deep learning with convolutional neural networks
复制标题

DOI:
10.1111/srt.12794
复制
发表时间:
2019-09-29
影响因子:
2.2
通讯作者:
Lee, Hwee Kuan
Lee, Hwee Kuan
中科院分区:
医学4区
文献类型:
--
作者:
Lim, Ziying Vanessa;Akram, Farhan;Lee, Hwee Kuan

文献摘要

被引文献

相似文献

医生对寻常痤疮的视觉评估和严重程度分级可能是主观的,导致观察者之间和观察者内部的差异。目的开发和验证一种用于自动计算研究者总体评估(伊加)量表的算法,以标准化痤疮严重程度和结果测量。材料与方法对416例痤疮患者的472张正面照片(检索时间:2004年1月1日-2017年4月8日)进行训练和测试。根据伊加量表,将照片标记为伊加清晰/几乎清晰(0-1)、伊加轻度(2)和伊加中度至重度(3-4)三组。分类模型使用卷积神经网络,模型分别在三种图像尺寸上进行训练。然后通过算法对照片进行分析,并将生成的自动伊加评分与临床评分进行比较。计算每个伊加等级标签的预测准确度和两个评分的一致性(Pearson相关性)。结果最佳分类准确率为67%。对于每个模型和各种图像输入大小,机器预测的分数和人类标签(临床评分和研究人员评分)之间的Pearson相关性为0.77。当在1200 x 1600的最大图像尺寸上使用Inception v4时,预测与临床评分的相关性最高。两组人类标签显示出0.77的高相关性,验证了地面真实标签的可重复性。混淆矩阵显示,模型在伊加2标签上的表现次优。利用高分辨率图像和大型数据集的深度学习技术将继续改进,显示出自动化临床图像分析和分级的潜力。
Background The visual assessment and severity grading of acne vulgaris by physicians can be subjective, resulting in inter- and intra-observer variability. Objective To develop and validate an algorithm for the automated calculation of the Investigator's Global Assessment (IGA) scale, to standardize acne severity and outcome measurements. Materials and Methods A total of 472 photographs (retrieved 01/01/2004-04/08/2017) in the frontal view from 416 acne patients were used for training and testing. Photographs were labeled according to the IGA scale in three groups of IGA clear/almost clear (0-1), IGA mild (2), and IGA moderate to severe (3-4). The classification model used a convolutional neural network, and models were separately trained on three image sizes. The photographs were then subjected to analysis by the algorithm, and the generated automated IGA scores were compared to clinical scoring. The prediction accuracy of each IGA grade label and the agreement (Pearson correlation) of the two scores were computed. Results The best classification accuracy was 67%. Pearson correlation between machine-predicted score and human labels (clinical scoring and researcher scoring) for each model and various image input sizes was 0.77. Correlation of predictions with clinical scores was highest when using Inception v4 on the largest image size of 1200 x 1600. Two sets of human labels showed a high correlation of 0.77, verifying the repeatability of the ground truth labels. Confusion matrices show that the models performed sub-optimally on the IGA 2 label. Conclusion Deep learning techniques harnessing high-resolution images and large datasets will continue to improve, demonstrating growing potential for automated clinical image analysis and grading.