A self-inspected adaptive SMOTE algorithm (SASMOTE) for highly imbalanced data classification in healthcare.

A self-inspected adaptive SMOTE algorithm (SASMOTE) for highly imbalanced data classification in healthcare.
复制标题

一种自检自适应 SMOTE 算法 (SASMOTE),用于医疗保健中高度不平衡的数据分类。

DOI:
10.1186/s13040-023-00330-4
复制
发表时间:
2023-04-25
期刊:
影响因子:
4.5
通讯作者:
--
中科院分区:
生物学3区
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

在许多医疗保健应用中,用于分类的数据集可能由于诸如疾病发作的目标事件的罕见发生而高度不平衡。SMOTE(Synthetic Minority Over-sampling Technique)算法是一种通过对少数类进行过采样来实现不平衡数据分类的有效方法。然而,由SMOTE生成的样本可能是模糊的、低质量的并且与多数类不可分离。为了提高生成样本的质量,我们提出了一种新的自检测自适应SMOTE(SASMOTE)模型,该模型利用自适应最近邻选择算法来识别“可见”最近邻,用于生成可能属于少数类的样本。为了进一步提高生成的样本的质量,通过自检方法的不确定性消除引入建议SASMOTE模型。它的目标是过滤掉生成的样本是高度不确定的,与多数类不可分割的。所提出的算法的有效性进行了比较,现有的基于SMOTE的算法,并证明通过两个现实世界的案例研究,在医疗保健,包括风险基因的发现和致命的先天性心脏病的预测。通过生成更高质量的合成样本,与其他方法相比,所提出的算法能够帮助实现更好的平均预测性能(就F1得分而言),这有望增强机器学习模型在高度不平衡的医疗数据上的可用性。
In many healthcare applications, datasets for classification may be highly imbalanced due to the rare occurrence of target events such as disease onset. The SMOTE (Synthetic Minority Over-sampling Technique) algorithm has been developed as an effective resampling method for imbalanced data classification by oversampling samples from the minority class. However, samples generated by SMOTE may be ambiguous, low-quality and non-separable with the majority class. To enhance the quality of generated samples, we proposed a novel self-inspected adaptive SMOTE (SASMOTE) model that leverages an adaptive nearest neighborhood selection algorithm to identify the “visible” nearest neighbors, which are used to generate samples likely to fall into the minority class. To further enhance the quality of the generated samples, an uncertainty elimination via self-inspection approach is introduced in the proposed SASMOTE model. Its objective is to filter out the generated samples that are highly uncertain and inseparable with the majority class. The effectiveness of the proposed algorithm is compared with existing SMOTE-based algorithms and demonstrated through two real-world case studies in healthcare, including risk gene discovery and fatal congenital heart disease prediction. By generating the higher quality synthetic samples, the proposed algorithm is able to help achieve better prediction performance (in terms of F1 score) on average compared to the other methods, which is promising to enhance the usability of machine learning models on highly imbalanced healthcare data.
DOI: 10.1186/1471-2288-14-137
发表时间: 2014-12-22
影响因子: 4
作者:
van der Ploeg T;Austin PC;Steyerberg EW
通讯作者: Steyerberg EW
DOI: 10.1016/j.neuroimage.2013.10.005
发表时间: 2014-02-15
期刊: NeuroImage
影响因子: 5.7
作者:
Dubey R;Zhou J;Wang Y;Thompson PM;Ye J;Alzheimer's Disease Neuroimaging Initiative
通讯作者: Alzheimer's Disease Neuroimaging Initiative
DOI: 10.1186/s13040-016-0117-1
发表时间: 2016
期刊: BioData mining
影响因子: 4.5
作者:
Li J;Fong S;Sung Y;Cho K;Wong R;Wong KKL
通讯作者: Wong KKL
DOI: 10.1109/tpami.2007.70735
发表时间: 2008-05-01
影响因子: 23.6
作者:
Lin, Tong;Zha, Hongbin
通讯作者: Zha, Hongbin
DOI: 10.1093/oxfordjournals.aje.a009302
发表时间: 1997-09-15
影响因子: 5
作者:
Chambless, LE;Heiss, G;Clegg, LX
通讯作者: Clegg, LX