COCOA: COrrelation COefficient-Aware Data Augmentation

COCOA: COrrelation COefficient-Aware Data Augmentation
复制标题

COCOA:相关系数感知数据增强

DOI:
--
复制
发表时间:
2021
期刊:
International Conference on Extending Database Technology
影响因子:
--
通讯作者:
Ziawasch Abedjan
Ziawasch Abedjan
中科院分区:
--
文献类型:
--
作者:
Mahdi Esmailoghli;Jorge;Ziawasch Abedjan

文献摘要

参考文献

被引文献

相似文献

计算相关系数是数据科学中最常用的测量方法之一。尽管线性相关性快速且易于计算,但它们在非线性关联存在的情况下缺乏鲁棒性和有效性。基于等级的系数(例如 Spearman 系数)更合适。然而,基于排名的度量首先需要对值进行排序并获得排名,从而使其计算超线性。受此影响的用例之一是通过从大型数据库中提取特征来丰富机器学习 (ML) 的数据。从数百万个候选特征中找到最有希望的特征以提高机器学习的准确性需要数十亿次相关计算。在本文中,我们引入了一种索引结构,可确保在线性时间内进行基于排名的相关性计算。我们的解决方案在数据丰富设置中将相关性计算速度提高了 500 倍。
Calculating correlation coefficients is one of the most used measures in data science. Although linear correlations are fast and easy to calculate, they lack robustness and effectiveness in the existence of non-linear associations. Rank-based coefficients such as Spearman’s are more suitable. However, rank-based measures first require to sort the values and obtain the ranks, making their calculation super-linear. One of the use-cases that is affected by this is data enrichment for Machine Learning (ML) through feature extraction from large databases. Finding the most promising features from millions of candidates to increase the ML accuracy requires billions of correlation calculations. In this paper, we introduce an index structure that ensures rank-based correlation calculation in a linear time. Our solution accelerates the correlation calculation up to 500 times in the data enrichment setting.
DOI: 10.1145/3318464.3389726
发表时间: 2020-06
期刊: Proceedings. ACM-SIGMOD International Conference on Management of Data
影响因子: --
作者:
Zhang Y;Ives ZG
通讯作者: Ives ZG