Does feature selection improve classification accuracy? Impact of sample size and feature selection on classification using anatomical magnetic resonance images

Does feature selection improve classification accuracy? Impact of sample size and feature selection on classification using anatomical magnetic resonance images
复制标题

DOI:
10.1016/j.neuroimage.2011.11.066
复制
发表时间:
2012-03-01
期刊:
影响因子:
5.7
通讯作者:
Lin, ChingPo
Lin, ChingPo
中科院分区:
医学1区
文献类型:
--
作者:
Chu, Carlton;Hsu, Ai-Ling;Lin, ChingPo

文献摘要

被引文献

相似文献

越来越多的研究使用机器学习方法来表征从神经成像数据中可辨别的解剖差异模式。图像数据的高维性常常引起人们的关注,即需要特征选择来获得最佳精度。在以前的研究中,大多使用固定的样本量,一些显示出更高的预测精度与特征选择,而其他人没有。在这项研究中,我们比较了四种常见的特征选择方法。1)基于先验知识的预先选择的感兴趣区域(ROI)。2)单变量t检验滤波。3)递归特征消除(RFE),以及4)受ROI约束的t检验滤波。从不同的样本量,有和没有特征选择,实现的预测精度进行了统计学比较。为了证明效果,我们使用从阿尔茨海默病神经成像倡议(ADNI)收集的T1加权解剖扫描中分割的灰质作为线性支持向量机分类器的输入特征。目的是描述阿尔茨海默病(AD)患者和认知正常受试者之间的差异模式,并描述轻度认知障碍(MCI)患者和正常受试者之间的差异。此外,我们还比较了在12个月内转换为AD的MCI患者和未转换的MCI患者之间的分类准确性。两种数据驱动的特征选择方法(t检验过滤和RFE)的预测准确性并不比使用全脑数据实现的准确性更好。我们发现,我们可以实现最准确的表征,通过使用预期神经变性(海马和海马旁回)的先验知识。因此,特征选择确实提高了分类精度,但这取决于所采用的方法。一般而言,较大的样本量产生较高的准确度,而通过使用现有文献中的知识获得的优势较小。(C)2011 Elsevier Inc. All rights reserved.
There are growing numbers of studies using machine learning approaches to characterize patterns of anatomical difference discernible from neuroimaging data. The high-dimensionality of image data often raises a concern that feature selection is needed to obtain optimal accuracy. Among previous studies, mostly using fixed sample sizes, some show greater predictive accuracies with feature selection, whereas others do not. In this study, we compared four common feature selection methods. 1) Pre-selected region of interests (ROIs) that are based on prior knowledge. 2) Univariate t-test filtering. 3) Recursive feature elimination (RFE), and 4) t-test filtering constrained by ROIs. The predictive accuracies achieved from different sample sizes, with and without feature selection, were compared statistically. To demonstrate the effect, we used grey matter segmented from the T1-weighted anatomical scans collected by the Alzheimer's disease Neuroimaging Initiative (ADNI) as the input features to a linear support vector machine classifier. The objective was to characterize the patterns of difference between Alzheimer's disease (AD) patients and cognitively normal subjects, and also to characterize the difference between mild cognitive impairment (MCI) patients and normal subjects. In addition, we also compared the classification accuracies between MCI patients who converted to AD and MCI patients who did not convert within the period of 12 months. Predictive accuracies from two data-driven feature selection methods (t-test filtering and RFE) were no better than those achieved using whole brain data. We showed that we could achieve the most accurate characterizations by using prior knowledge of where to expect neurodegeneration (hippocampus and parahippocampal gyrus). Therefore, feature selection does improve the classification accuracies, but it depends on the method adopted. In general, larger sample sizes yielded higher accuracies with less advantage obtained by using knowledge from the existing literature. (C) 2011 Elsevier Inc. All rights reserved.