Wise Feature Selection for Breast Cancer Detection from a Clinical Dataset

Wise Feature Selection for Breast Cancer Detection from a Clinical Dataset
复制标题

从临床数据集中进行乳腺癌检测的明智特征选择

DOI:
10.1109/icbme54433.2021.9750287
复制
发表时间:
2021
期刊:
Iranian Conference on Biomedical Engineering
影响因子:
--
通讯作者:
M. Vali
M. Vali
中科院分区:
--
文献类型:
--
作者:
Mahsa Bahrami;M. Vali

文献摘要

被引文献

相似文献

乳腺癌是一种常见的癌症,尤其是女性。乳腺癌的早期发现被认为是一个额外的重要因素,因为它降低了死亡率并便于治疗。需要用于检测乳腺癌的精确自动算法。在本文中,我们开发了一些不同的特征选择方法,用于准确的乳腺癌检测。对临床数据进行预处理,然后基于特征重要性进行明智的特征选择。除此之外,主成分分析(PCA),增量PCA,核PCA,独立成分分析,因子分析和奇异值分解方法的实施和分析的特征选择和降维。最后采用多层感知器进行分类。在具有569个记录的乳腺癌威斯康星州诊断数据集上评估特征选择方法的性能。使用特征重要性方法,测试数据的最佳准确性、敏感性、特异性、F1评分和Cohen's kappa分别为97.4%、98.6%、95.3%、97.6%和0.94。
Breast cancer is a common cancer, especially in women. Early detection of breast cancer is taken an added importance because it alleviates the rate of mortality and facilitates treatment. Accurate automatic algorithms for the detection of breast cancer are needed. In this paper, we developed a number of different feature selection methods for accurate breast cancer detection. Clinical data was pre-processed and then wise feature selection based on feature importance was applied for feature selection. In addition to this, principal component analysis (PCA), incremental PCA, kernel PCA, independent component analysis, factor analysis, and singular value decomposition methods were implemented and analyzed for feature selection and dimension reduction. Finally, multi-layer perceptron was used for classification. The performance of feature selection methods was evaluated on Breast Cancer Wisconsin Diagnostic dataset with 569 recordings. The best accuracy, sensitivity, specificity, F1-score, and Cohen's kappa on the test data were 97.4%, 98.6%, 95.3%, 97.6%, and 0.94 respectively, with a feature importance method.