Foundations of data imbalance and solutions for a data democracy

Foundations of data imbalance and solutions for a data democracy
复制标题

DOI:
10.1016/b978-0-12-818366-3.00005-8
复制
发表时间:
2020-01-01
期刊:
DATA DEMOCRACY: AT THE NEXUS OF ARTIFICIAL INTELLIGENCE, SOFTWARE DEVELOPMENT, AND KNOWLEDGE ENGINEERING
影响因子:
--
通讯作者:
Batarseh, Feras A.
Batarseh, Feras A.
中科院分区:
其他
文献类型:
--
作者:
Kulkarni, Ajay;Chong, Deri;Batarseh, Feras A.

文献摘要

被引文献

相似文献

在对数据集进行分类时,处理不平衡数据是一个普遍存在的问题。很多时候,这个问题会在决策或实施政策时造成偏见。因此,了解导致数据不平衡(或类不平衡)的因素至关重要。这种隐藏的偏见和不平衡可能导致数据暴政,并对数据民主构成重大挑战。在本章中,解决了两个重要的统计要素:阶级不平衡程度和概念的复杂性;解决这些问题有助于建立数据民主的基础。此外,还讨论了适用于这些场景的统计度量,并在现实数据集(汽车保险索赔)上实现了这些度量。最后,在Python中实现了流行的数据级方法,如随机过采样、随机欠采样、合成少数派过采样技术、Tomek link等,并比较了它们的性能。
Dealing with imbalanced data is a prevalent problem while performing classification on the datasets. Many times, this problem contributes to bias while making decisions or implementing policies. Thus, it is vital to understand the factors which cause imbalance in the data (or class imbalance). Such hidden biases and imbalances can lead to data tyranny and a major challenge to a data democracy. In this chapter, two essential statistical elements are resolved: the degree of class imbalance and the complexity of the concept; solving such issues helps in building the foundations of a data democracy. Furthermore, statistical measures which are appropriate in these scenarios are discussed and implemented on a real-life dataset (car insurance claims). In the end, popular data-level methods such as random oversampling, random undersampling, synthetic minority oversampling technique, Tomek link, and others are implemented in Python, and their performance is compared.