Performance evaluation of deep neural ensembles toward malaria parasite detection in thin-blood smear images

Performance evaluation of deep neural ensembles toward malaria parasite detection in thin-blood smear images
复制标题

DOI:
10.7717/peerj.6977
复制
发表时间:
2019-05-28
期刊:
影响因子:
2.7
通讯作者:
Antani, Sameer K.
Antani, Sameer K.
中科院分区:
生物学3区
文献类型:
--
作者:
Rajaraman, Sivaramakrishnan;Jaeger, Stefan;Antani, Sameer K.

文献摘要

被引文献

相似文献

背景。疟疾是一种威胁生命的疾病,由感染红细胞(红细胞)的疟原虫引起。显微镜厚/薄膜血检查中寄生细胞的人工鉴定和计数仍然是常见的疾病诊断方法,但方法繁琐。其诊断准确性受到观察者之间/内部可变性的不利影响,特别是在资源有限的情况下进行大规模筛查。基于数据驱动的深度学习算法(如卷积神经网络(CNN))的最先进的计算机辅助诊断工具已经成为图像识别任务的首选架构。然而,由于cnn对训练数据波动的敏感性,其方差较大,可能会出现过拟合。本研究的主要目的是通过构建模型集合来检测薄血涂片图像中的寄生细胞,从而减少模型方差,提高鲁棒性和泛化性。我们评估了自定义和预训练cnn的性能,并构建了一个最优模型集合,以应对薄血涂片图像中寄生细胞和正常细胞的分类挑战。交叉验证研究在患者水平上进行,以确保防止数据泄露到验证中并减少泛化错误。根据下列业绩指标评价这些模型:(a)准确性;(b)受试者工作特征曲线下面积(AUC);(c)均方误差(MSE);(d)精度;(e) f值;(f) Matthews相关系数(MCC)。我们观察到,使用VGG-19和SqueezeNet构建的集成模型在对寄生和未感染细胞进行分类的几个性能指标上优于最新的技术,有助于改进疾病筛查。集成学习通过优化组合多个模型的预测来减少模型方差,降低对训练数据和训练算法选择的敏感性。模型集成的性能模拟了真实世界的条件,减少了方差,过度拟合,并导致改进的泛化。
Background. Malaria is a life-threatening disease caused by Plasmodium parasites that infect the red blood cells (RBCs). Manual identification and counting of parasitized cells in microscopic thick/thin-film blood examination remains the common, but burdensome method for disease diagnosis. Its diagnostic accuracy is adversely impacted by inter/intra-observer variability, particularly in large-scale screening under resource-constrained settings.Introduction. State-of-the-art computer-aided diagnostic tools based on data-driven deep learning algorithms like convolutional neural network (CNN) has become the architecture of choice for image recognition tasks. However, CNNs suffer from high variance and may overfit due to their sensitivity to training data fluctuations.Objective. The primary aim of this study is to reduce model variance, improve robustness and generalization through constructing model ensembles toward detecting parasitized cells in thin-blood smear images.Methods. We evaluate the performance of custom and pretrained CNNs and construct an optimal model ensemble toward the challenge of classifying parasitized and normal cells in thin-blood smear images. Cross-validation studies are performed at the patient level to ensure preventing data leakage into the validation and reduce generalization errors. The models are evaluated in terms of the following performance metrics: (a) Accuracy; (b) Area under the receiver operating characteristic (ROC) curve (AUC); (c) Mean squared error (MSE); (d) Precision; (e) F-score; and (f) Matthews Correlation Coefficient (MCC).Results. It is observed that the ensemble model constructed with VGG-19 and SqueezeNet outperformed the state-of-the-art in several performance metrics toward classifying the parasitized and uninfected cells to aid in improved disease screening.Conclusions. Ensemble learning reduces the model variance by optimally combining the predictions of multiple models and decreases the sensitivity to the specifics of training data and selection of training algorithms. The performance of the model ensemble simulates real-world conditions with reduced variance, overfitting and leads to improved generalization.