MABAL: a Novel Deep-Learning Architecture for Machine-Assisted Bone Age Labeling

MABAL: a Novel Deep-Learning Architecture for Machine-Assisted Bone Age Labeling
复制标题

DOI:
10.1007/s10278-018-0053-3
复制
发表时间:
2018-08-01
影响因子:
4.4
通讯作者:
Ayyala, Rama
Ayyala, Rama
中科院分区:
工程技术2区
文献类型:
--
作者:
Mutasa, Simukayi;Chang, Peter D.;Ayyala, Rama

文献摘要

被引文献

相似文献

骨龄评估(BAA)是儿科放射学中常用的诊断研究,用于评估骨骼成熟度。评估BAA最常用的方法是Greulich和Pyle方法(Pediatr Radiol 46.9:1269-1274,2016; Arch Dis Child 81.2:172-173,1999)图谱。BAA的评价对于放射科医生来说可能是一个繁琐且耗时的过程。因此,已经提出了几种计算机辅助检测/诊断(CAD)方法用于BAA的自动化。传统的CAD工具传统上依赖于硬编码的算法功能的BAA遭受各种各样的缺点。最近,卷积神经网络(CNN)的出现和扩散在各种医学成像应用中显示出了希望。至少有两个已发表的使用深度学习评估骨龄的应用(Med Image Anal 36:41-51,2017; JDI 1-5,2017)。然而,当前的实现受到架构设计和相对较小的数据集的组合的限制。本研究的目的是证明定制神经网络算法的好处,该算法经过仔细校准,以利用相对较大的机构数据集评估骨龄。在这样做的过程中,本研究的目的是表明,先进的架构可以在医学成像领域从头开始成功训练,并可以生成优于任何现有算法的结果。训练数据由10,289张不同骨龄检查的图像组成,8909来自我们机构的医院图像采集和通信系统,1383来自公共数字手图谱数据库。这些数据被分为四组,8岁以上的男女儿童各一组,10岁以下的男女儿童各一组。测试集包括每个1岁年龄组(0 - 1岁至14-15岁+)的20张X线片,男性和女性各占一半。测试集包括用于骨龄评估的左手X线片、无显著发现的创伤评估和骨骼调查。为这项研究设计了一个14隐层定制的神经网络。该网络包括几种最先进的技术,包括残差式连接、初始层和空间Transformer层。对网络输入应用数据增强以防止过拟合。使用线性回归输出。均方误差被用作网络损失函数,平均绝对误差(MAE)被用作主要的性能指标。年轻女性验证集和测试集的MAE准确率分别为0.654和0.561。对于老年女性,验证和测试精度分别为0.662和0.497。对于年轻男性,验证和测试精度分别为0.649和0.585。最后,对于老年男性,验证和测试集的准确度分别为0.581和0.501。女性队列每个训练900个时期,男性队列训练600个时期。采用八重交叉验证集进行超参数调整。测试误差是在使用选定的超参数对完整数据集进行训练后获得的。使用我们提出的定制神经网络架构对我们的大量可用数据,我们实现了聚合验证和测试集平均绝对误差分别为0.637和0.536。迄今为止,这是利用深度学习进行骨龄评估的最佳表现。我们的研究结果支持了我们最初的假设,即定制的专用神经网络比来自预先训练的成像数据集的网络提供了更好的性能。我们在最初的工作基础上,通过添加最先进的技术,如残差连接和初始架构,进一步提高了预测精度。这一点很重要,因为目前使用残差和/或初始架构的假设是,考虑到医学成像中相对较小的数据集,成功实现需要大型预训练网络。相反,我们证明了一个包含先进CNN策略的小型定制架构确实可以从头开始训练,从而显著提高算法的准确性。应该注意的是,对于所有四个队列,检验误差优于验证误差。其中一个原因是,我们的测试集的基础事实是通过对两个儿科放射科医生读数进行平均来获得的,而我们的训练数据只使用了一个读数。这表明,尽管训练数据相对嘈杂,但该算法可以成功地对观察者之间的变化进行建模,并生成接近预期地面真实的估计值。
Bone age assessment (BAA) is a commonly performed diagnostic study in pediatric radiology to assess skeletal maturity. The most commonly utilized method for assessment of BAA is the Greulich and Pyle method (Pediatr Radiol 46.9:1269-1274, 2016; Arch Dis Child 81.2:172-173, 1999) atlas. The evaluation of BAA can be a tedious and time-consuming process for the radiologist. As such, several computer-assisted detection/diagnosis (CAD) methods have been proposed for automation of BAA. Classical CAD tools have traditionally relied on hard-coded algorithmic features for BAA which suffer from a variety of drawbacks. Recently, the advent and proliferation of convolutional neural networks (CNNs) has shown promise in a variety of medical imaging applications. There have been at least two published applications of using deep learning for evaluation of bone age (Med Image Anal 36:41-51, 2017; JDI 1-5, 2017). However, current implementations are limited by a combination of both architecture design and relatively small datasets. The purpose of this study is to demonstrate the benefits of a customized neural network algorithm carefully calibrated to the evaluation of bone age utilizing a relatively large institutional dataset. In doing so, this study will aim to show that advanced architectures can be successfully trained from scratch in the medical imaging domain and can generate results that outperform any existing proposed algorithm.The training data consisted of 10,289 images of different skeletal age examinations, 8909 from the hospital Picture Archiving and Communication System at our institution and 1383 from the public Digital Hand Atlas Database. The data was separated into four cohorts, one each for male and female children above the age of 8, and one each for male and female children below the age of 10. The testing set consisted of 20 radiographs of each 1-year-age cohort from 0 to 1 years to 14-15+ years, half male and half female. The testing set included left-hand radiographs done for bone age assessment, trauma evaluation without significant findings, and skeletal surveys. A 14 hidden layer-customized neural network was designed for this study. The network included several state of the art techniques including residual-style connections, inception layers, and spatial transformer layers. Data augmentation was applied to the network inputs to prevent overfitting. A linear regression output was utilized. Mean square error was used as the network loss function and mean absolute error (MAE) was utilized as the primary performance metric. MAE accuracies on the validation and test sets for young females were 0.654 and 0.561 respectively. For older females, validation and test accuracies were 0.662 and 0.497 respectively. For young males, validation and test accuracies were 0.649 and 0.585 respectively. Finally, for older males, validation and test set accuracies were 0.581 and 0.501 respectively. The female cohorts were trained for 900 epochs each and the male cohorts were trained for 600 epochs. An eightfold cross-validation set was employed for hyperparameter tuning. Test error was obtained after training on a full data set with the selected hyperparameters. Using our proposed customized neural network architecture on our large available data, we achieved an aggregate validation and test set mean absolute errors of 0.637 and 0.536 respectively. To date, this is the best published performance on utilizing deep learning for bone age assessment. Our results support our initial hypothesis that customized, purpose-built neural networks provide improved performance over networks derived from pre-trained imaging data sets. We build on that initial work by showing that the addition of state-of-the-art techniques such as residual connections and inception architecture further improves prediction accuracy. This is important because the current assumption for use of residual and/or inception architectures is that a large pre-trained network is required for successful implementation given the relatively small datasets in medical imaging. Instead we show that a small, customized architecture incorporating advanced CNN strategies can indeed be trained from scratch, yielding significant improvements in algorithm accuracy. It should be noted that for all four cohorts, testing error outperformed validation error. One reason for this is that our ground truth for our test set was obtained by averaging two pediatric radiologist reads compared to our training data for which only a single read was used. This suggests that despite relatively noisy training data, the algorithm could successfully model the variation between observers and generate estimates that are close to the expected ground truth.