Analysis of sampling techniques for imbalanced data: An n = 648 ADNI study.

Analysis of sampling techniques for imbalanced data: An n = 648 ADNI study.
复制标题

DOI:
10.1016/j.neuroimage.2013.10.005
复制
发表时间:
2014-02-15
期刊:
影响因子:
5.7
通讯作者:
Alzheimer's Disease Neuroimaging Initiative
Alzheimer's Disease Neuroimaging Initiative
中科院分区:
医学1区
文献类型:
--
作者:
Dubey R;Zhou J;Wang Y;Thompson PM;Ye J;Alzheimer's Disease Neuroimaging Initiative

文献摘要

参考文献

被引文献

相似文献

许多神经成像应用处理不平衡的成像数据。例如,在阿尔茨海默病神经成像倡议(ADNI)数据集中,适合于该研究的轻度认知障碍(MCI)病例是阿尔茨海默病(AD)患者的结构磁共振成像(MRI)模式的近两倍,是蛋白质组学模式的对照病例的六倍。从不平衡数据中构造一个准确的分类器是一项具有挑战性的任务。传统的分类器,旨在最大限度地提高整体的预测精度往往会将所有的数据分类到大多数类。在本文中,我们研究了一个集成系统的特征选择和数据采样的类不平衡问题。我们系统地分析了各种采样技术,通过检查不同的比率和类型的欠采样,过采样,以及过采样和欠采样方法的组合的有效性。我们彻底研究了六种广泛使用的特征选择算法,以识别重要的生物标志物,从而降低数据的复杂性。使用两种不同的分类器,包括随机森林和支持向量机的基础上的分类精度,在接收器工作特征曲线(AUC),灵敏度和特异性措施的面积集成技术的疗效进行评估。我们广泛的实验结果表明,对于ADNI中的各种问题设置,(1)。在不同的数据采样技术和无采样方法中,使用基于K-中心点技术的欠采样获得的平衡训练集给出了最佳的整体性能;以及(2).具有稳定性选择的稀疏逻辑回归在各种特征选择算法中具有竞争力的性能。各种设置的综合实验表明,我们提出的多个欠采样数据集的集成模型产生稳定和有前途的结果。
Many neuroimaging applications deal with imbalanced imaging data. For example, in Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset, the mild cognitive impairment (MCI) cases eligible for the study are nearly two times the Alzheimer’s disease (AD) patients for structural magnetic resonance imaging (MRI) modality and six times the control cases for proteomics modality. Constructing an accurate classifier from imbalanced data is a challenging task. Traditional classifiers that aim to maximize the overall prediction accuracy tend to classify all data into the majority class. In this paper, we study an ensemble system of feature selection and data sampling for the class imbalance problem. We systematically analyze various sampling techniques by examining the efficacy of different rates and types of undersampling, oversampling, and a combination of over and under sampling approaches. We thoroughly examine six widely used feature selection algorithms to identify significant biomarkers and thereby reduce the complexity of the data. The efficacy of the ensemble techniques is evaluated using two different classifiers including Random Forest and Support Vector Machines based on classification accuracy, area under the receiver operating characteristic curve (AUC), sensitivity, and specificity measures. Our extensive experimental results show that for various problem settings in ADNI, (1). a balanced training set obtained with K-Medoids technique based undersampling gives the best overall performance among different data sampling techniques and no sampling approach; and (2). sparse logistic regression with stability selection achieves competitive performance among various feature selection algorithms. Comprehensive experiments with various settings show that our proposed ensemble model of multiple undersampled datasets yields stable and promising results.
DOI: 10.1016/s0197-4580(01)00271-8
发表时间: 2001-09-01
影响因子: 4.2
作者:
Dickerson, BC;Goncharova, I;deToledo-Morrell, L
通讯作者: deToledo-Morrell, L
DOI: 10.1038/nrneurol.2009.215
发表时间: 2010-02
影响因子: 38.1
作者:
Frisoni, Giovanni B.;Fox, Nick C.;Jack, Clifford R., Jr.;Scheltens, Philip;Thompson, Paul M.
通讯作者: Thompson, Paul M.
DOI: 10.1212/01.wnl.0000256697.20968.d7
发表时间: 2007-03-13
期刊: NEUROLOGY
影响因子: 9.9
作者:
Devanand, D. P.;Pradhaban, G.;de Leon, M. J.
通讯作者: de Leon, M. J.
DOI: 10.1613/jair.953
发表时间: 2002-01-01
影响因子: 5
作者:
Chawla, NV;Bowyer, KW;Kegelmeyer, WP
通讯作者: Kegelmeyer, WP
DOI: 10.1111/j.0824-7935.2004.t01-1-00228.x
发表时间: 2004-02-01
影响因子: 2.8
作者:
Estabrooks, A;Jo, TH;Japkowicz, N
通讯作者: Japkowicz, N