L2M NSERC - Intelligent system for classifying imbalanced data based on three-way Bayesian confirmation
L2M NSERC - Intelligent system for classifying imbalanced data based on three-way Bayesian confirmation
批准号:
580671-2023
负责人:
Yao, YiyuYY
金额:
$1.46万
依托单位:
依托单位国家:
加拿大
项目类别:
Idea to Innovation
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Imbalanced data refers to a data set where the data points are not evenly distributed across different classes. As a result, there are majority classes taking high proportions of the data points and minority classes taking the remaining low proportions. Imbalanced data is prevalent in practical situations, especially when we try to detect something abnormal such as fraudulent transactions, spam emails, and certain diseases. Standard classification models may not work well with imbalanced data. For example, consider an imbalanced data set where a majority class takes 90% of the data points and a minority class takes 10%. A standard classification model will likely learn that it can simply predict the majority class without any condition and be correct for 90% of the cases. This definitely does not explain the inherent reason for the classifications. In contrast, it is hard to predict a minority class correctly in most cases. However, the minority class usually consists of fraudulent transactions, spam emails, and cases of diseases that we actually want to learn and detect, more importantly than the majority class.We apply the Bayesian confirmation theory to build an intelligent system to learn effective features and rules for classifying imbalanced data. The importance of a certain feature or a group of features is evaluated by comparing the prior probability of a class before observing the value(s) of the feature(s) and the posterior probability after observing the value(s). The change shows the real impact of the feature(s) on our classification decisions and helps detect the effective features. For example, in detecting a certain disease, we may have a group of features measured through different tests. The values of some features may significantly increase or decrease the probability of the disease. Accordingly, the patient may be asked to do the corresponding tests to achieve an effective diagnosis. Compared to the standard classification models, we are able to learn the actual effective features and rules for detecting the majority and minority classes, addressing the aforementioned issues in analyzing imbalanced data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金