L2M NSERC - Intelligent system for classifying imbalanced data based on three-way Bayesian confirmation
L2M NSERC - Intelligent system for classifying imbalanced data based on three-way Bayesian confirmation
批准号:
580671-2023
负责人:
Yao, YiyuYY
金额:
$1.46万
依托单位:
依托单位国家:
加拿大
项目类别:
Idea to Innovation
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
不平衡数据是指数据点不均匀分布在不同类中的数据集。因此,多数类占据了高比例的数据点,少数类占据了低比例的数据点。不平衡数据在实际情况中很普遍,特别是当我们试图检测一些异常情况时,例如欺诈性交易、垃圾邮件和某些疾病。标准分类模型可能不能很好地处理不平衡数据。例如,考虑一个不平衡的数据集,其中多数类占90%的数据点,少数类占10%。一个标准的分类模型可能会学习到,它可以在没有任何条件的情况下简单地预测大多数类别,并且在90%的情况下是正确的。这绝对不能解释分类的内在原因。相比之下,在大多数情况下,很难正确预测一个少数阶级。然而,少数类通常包括欺诈性交易、垃圾邮件,以及我们实际上想要学习和检测的疾病案例,这些比大多数类更重要。我们运用贝叶斯确认理论建立了一个智能系统来学习对不平衡数据进行分类的有效特征和规则。通过比较观察某一类特征值之前的先验概率和观察该特征值之后的后验概率来评估某一类特征或一组特征的重要性。这个变化显示了特征对我们分类决策的真实影响,并有助于检测有效的特征。例如,在检测某种疾病时,我们可能会通过不同的测试来测量一组特征。某些特征的值可能显著增加或减少疾病的可能性。因此,可以要求患者做相应的检查,以获得有效的诊断。与标准的分类模型相比,我们能够学习到实际的有效特征和规则来检测多数类和少数类,解决了上述在分析不平衡数据时的问题。
英文摘要
Imbalanced data refers to a data set where the data points are not evenly distributed across different classes. As a result, there are majority classes taking high proportions of the data points and minority classes taking the remaining low proportions. Imbalanced data is prevalent in practical situations, especially when we try to detect something abnormal such as fraudulent transactions, spam emails, and certain diseases. Standard classification models may not work well with imbalanced data. For example, consider an imbalanced data set where a majority class takes 90% of the data points and a minority class takes 10%. A standard classification model will likely learn that it can simply predict the majority class without any condition and be correct for 90% of the cases. This definitely does not explain the inherent reason for the classifications. In contrast, it is hard to predict a minority class correctly in most cases. However, the minority class usually consists of fraudulent transactions, spam emails, and cases of diseases that we actually want to learn and detect, more importantly than the majority class.We apply the Bayesian confirmation theory to build an intelligent system to learn effective features and rules for classifying imbalanced data. The importance of a certain feature or a group of features is evaluated by comparing the prior probability of a class before observing the value(s) of the feature(s) and the posterior probability after observing the value(s). The change shows the real impact of the feature(s) on our classification decisions and helps detect the effective features. For example, in detecting a certain disease, we may have a group of features measured through different tests. The values of some features may significantly increase or decrease the probability of the disease. Accordingly, the patient may be asked to do the corresponding tests to achieve an effective diagnosis. Compared to the standard classification models, we are able to learn the actual effective features and rules for detecting the majority and minority classes, addressing the aforementioned issues in analyzing imbalanced data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金