Standardized evaluation of algorithms for computer-aided diagnosis of dementia based on structural MRI: the CADDementia challenge.

Standardized evaluation of algorithms for computer-aided diagnosis of dementia based on structural MRI: the CADDementia challenge.
复制标题

DOI:
10.1016/j.neuroimage.2015.01.048
复制
发表时间:
2015-05-01
期刊:
影响因子:
5.7
通讯作者:
Alzheimer's Disease Neuroimaging Initiative
Alzheimer's Disease Neuroimaging Initiative
中科院分区:
医学1区
文献类型:
--
作者:
Bron EE;Smits M;van der Flier WM;Vrenken H;Barkhof F;Scheltens P;Papma JM;Steketee RM;Méndez Orellana C;Meijboom R;Pinto M;Meireles JR;Garrett C;Bastos-Leite AJ;Abdulkadir A;Ronneberger O;Amoroso N;Bellotti R;Cárdenas-Peña D;Álvarez-Meza AM;Dolph CV;Iftekharuddin KM;Eskildsen SF;Coupé P;Fonov VS;Franke K;Gaser C;Ledig C;Guerrero R;Tong T;Gray KR;Moradi E;Tohka J;Routier A;Durrleman S;Sarica A;Di Fatta G;Sensi F;Chincarini A;Smith GM;Stoyanov ZV;Sørensen L;Nielsen M;Tangaro S;Inglese P;Wachinger C;Reuter M;van Swieten JC;Niessen WJ;Klein S;Alzheimer's Disease Neuroimaging Initiative

文献摘要

被引文献

相似文献

基于结构MRI的痴呆计算机辅助诊断算法在文献中表现出很高的性能,但由于使用不同的数据集和方法进行评估,因此难以进行比较。此外,目前还不清楚算法将如何处理以前未见过的数据,因此,当没有真实的机会使算法适应手头的数据时,它们将如何在临床实践中执行。为了解决这些可比性、可推广性和临床适用性问题,我们组织了一项重大挑战,旨在客观地比较基于临床代表性多中心数据集的算法。以临床实践为出发点,目标是再现临床诊断。因此,我们评估了三个诊断组的多类分类算法:可能患有阿尔茨海默病的患者,轻度认知障碍患者和健康对照组。基于临床标准的诊断被用作参考标准,因为它是最好的参考,尽管它有已知的局限性。为了进行评价,使用了之前未见过的测试集,包括354个T1加权MRI扫描,诊断设盲。15个研究团队参与了总共29种算法。在小训练集(n=30)上以及可选地在来自其他来源的数据(例如,阿尔茨海默病神经影像学倡议,澳大利亚成像生物标志物和生活方式的旗舰研究老化)。性能最好的算法产生了63.0%的准确性和78.8%的受试者工作特征曲线下面积(AUC)。一般来说,最好的表现是使用基于体素的形态测量学的特征提取或包括体积,皮质厚度,形状和强度的特征的组合。该挑战是开放的,通过基于网络的框架:http://caddementia.grand-challenge.org新的提交。
Algorithms for computer-aided diagnosis of dementia based on structural MRI have demonstrated high performance in the literature, but are difficult to compare as different data sets and methodology were used for evaluation. In addition, it is unclear how the algorithms would perform on previously unseen data, and thus, how they would perform in clinical practice when there is no real opportunity to adapt the algorithm to the data at hand. To address these comparability, generalizability and clinical applicability issues, we organized a grand challenge that aimed to objectively compare algorithms based on a clinically representative multi-center data set. Using clinical practice as starting point, the goal was to reproduce the clinical diagnosis. Therefore, we evaluated algorithms for multi-class classification of three diagnostic groups: patients with probable Alzheimer’s disease, patients with mild cognitive impairment and healthy controls. The diagnosis based on clinical criteria was used as reference standard, as it was the best available reference despite its known limitations. For evaluation, a previously unseen test set was used consisting of 354 T1-weighted MRI scans with the diagnoses blinded. Fifteen research teams participated with in total 29 algorithms. The algorithms were trained on a small training set (n=30) and optionally on data from other sources (e.g., the Alzheimer’s Disease Neuroimaging Initiative, the Australian Imaging Biomarkers and Lifestyle flagship study of aging). The best performing algorithm yielded an accuracy of 63.0% and an area under the receiver-operating-characteristic curve (AUC) of 78.8%. In general, the best performances were achieved using feature extraction based on voxel-based morphometry or a combination of features that included volume, cortical thickness, shape and intensity. The challenge is open for new submissions via the web-based framework: http://caddementia.grand-challenge.org.