Addressing the missing data challenge in multi-modal datasets for the diagnosis of Alzheimer’s disease

Addressing the missing data challenge in multi-modal datasets for the diagnosis of Alzheimer’s disease
复制标题

DOI:
10.1016/j.jneumeth.2022.109582
复制
发表时间:
2022-03
影响因子:
3
通讯作者:
M. Aghili;Solale Tabarestani;M. Adjouadi
M. Aghili;Solale Tabarestani;M. Adjouadi
中科院分区:
医学4区
文献类型:
--
作者:
M. Aghili;Solale Tabarestani;M. Adjouadi

文献摘要

相似文献

背景除了识别定义其早期发病的微妙变化之外,阿尔茨海默病的准确诊断和预后面临的挑战之一是缺乏足够的数据,加上数据缺失的挑战。尽管阿尔茨海默病神经影像倡议(ADNI)数据库中有许多参与者,但许多观察结果有很多缺失的特征,这往往导致许多正在进行的实验中排除潜在有价值的数据点,特别是在纵向研究中。新方法出于检查所有参与者的必要性,即使是那些缺少测试或成像模式的参与者,这项研究引起了人们对梯度提升(GB)算法的关注,该算法具有解决缺失值的固有能力。考虑的四组包括:认知正常 (CN)、早期轻度认知障碍 (EMCI)、晚期轻度认知障碍 (LMCI) 和阿尔茨海默病 (AD)。在应用支持向量机 (SVM) 和随机森林 (RF) 等最先进的分类器之前,已经研究了使用数值技术在通用数据集中插补(即替换)数据的影响,并与 GB 算法进行了比较。实证评估表明,与 SVM 和 RF 算法相比,GB 性能对缺失值具有高度弹性。然而,当与更复杂的插补技术(例如假设数据不完整性程度较低的软插补或 K 最近邻 (KNN) 算法)相结合时,这些算法可以得到改进。结果当在模型生成和测试阶段考虑所有样本(包括不完整样本)时,所有四类受试者的多类分类中的分类精度提高了 3%。与现有方法相比,与其他方法不同,所提出的方法解决了 ADNI 数据集的具有挑战性的多类分类问题存在不同程度的缺失数据点。它还提供了现有插补技术对逐块缺失数据的影响的比较研究。所提出方法的结果根据用于 AD 分类的黄金标准方法进行了验证。
BackgroundOne of the challenges facing accurate diagnosis and prognosis of Alzheimer’s disease, beyond identifying the subtle changes that define its early onset, is the scarcity of sufficient data compounded by the missing data challenge. Although there are many participants in the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database, many of the observations have a lot of missing features which often leads to the exclusion of potentially valuable data points in many ongoing experiments, especially in longitudinal studies.New methodsMotivated by the necessity of examining all participants, even those with missing tests or imaging modalities, this study draws attention to the Gradient Boosting (GB) algorithm which has an inherent capability of addressing missing values. The four groups considered include: Cognitively Normal (CN), Early Mild Cognitive Impairment (EMCI), Late Mild Cognitive Impairment (LMCI) and Alzheimer's Disease (AD). Prior to applying state of the art classifiers such as Support Vector Machine (SVM) and Random Forest (RF), the impact of imputing (i.e., replacing) data in common datasets with numerical techniques has been investigated and compared with the GB algorithm. Empirical evaluations show that the GB performance is highly resilient to missing values in comparison to SVM and RF algorithms. These latter algorithms can however be improved when coupled with more sophisticated imputation technique such as soft-impute or K-Nearest Neighbors (KNN) algorithm assuming low extent of data incompleteness.ResultsThe classification accuracy has been improved by up to 3% in the multiclass classification of all four classes of subjects when all the samples including the incomplete ones are considered during the model generation and testing phases.Comparison with existing methodsUnlike other methods, the proposed approach addresses the challenging multiclass classification of the ADNI dataset in the presence of different levels of missing data points. It also provides a comparative study on effects of existing imputation techniques on a block-wise missing data. Results of the proposed method are validated against gold standard methods used for AD classification.