What is the Value of Data? on Mathematical Methods for Data Quality Estimation

What is the Value of Data? on Mathematical Methods for Data Quality Estimation
复制标题

DOI:
10.1109/isit44484.2020.9174311
复制
发表时间:
2020-01
期刊:
2020 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
通讯作者:
Netanel Raviv;Siddhartha Jain;Jehoshua Bruck
Netanel Raviv;Siddhartha Jain;Jehoshua Bruck
中科院分区:
其他
文献类型:
--
作者:
Netanel Raviv;Siddhartha Jain;Jehoshua Bruck

文献摘要

相似文献

数据是信息时代最重要的资产之一,其社会影响无可争议。然而,缺乏评估数据质量的严格方法。在本文中,我们提出了一个正式的定义,一个给定的数据集的质量。我们通过我们称之为预期直径的数量来评估数据集的质量,该数量衡量了两个随机选择的解释它的假设之间的预期分歧,并且最近在主动学习中找到了应用。我们专注于布尔超平面,并利用傅立叶分析,代数和概率方法的集合来提出理论保证和实际解决方案,用于计算预期的直径。我们还研究了代数结构数据集上的预期直径的行为,进行实验,验证这种质量的概念,并证明了我们的技术的可行性。
Data is one of the most important assets of the information age, and its societal impact is undisputed. Yet, rigorous methods of assessing the quality of data are lacking. In this paper, we propose a formal definition for the quality of a given dataset. We assess a dataset’s quality by a quantity we call the expected diameter, which measures the expected disagreement between two randomly chosen hypotheses that explain it, and has recently found applications in active learning. We focus on Boolean hyperplanes, and utilize a collection of Fourier analytic, algebraic, and probabilistic methods to come up with theoretical guarantees and practical solutions for the computation of the expected diameter. We also study the behaviour of the expected diameter on algebraically structured datasets, conduct experiments that validate this notion of quality, and demonstrate the feasibility of our techniques.