ICARUS: Minimizing Human Effort in Iterative Data Completion

ICARUS: Minimizing Human Effort in Iterative Data Completion
复制标题

ICARUS:在迭代数据完成中最大限度地减少人类工作量

DOI:
--
复制
发表时间:
2018
影响因子:
2.5
通讯作者:
Arnab Nandi
Arnab Nandi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Protiva Rahman;C. Hebert;Arnab Nandi

文献摘要

被引文献

相似文献

数据准备的一个重要步骤是处理不完整的数据集。在某些情况下,缺失值未被报告,因为它们是域的特征,并且为从业者所知。由于缺失值的这种性质,插补和推理方法不起作用,需要领域专家的输入。专家填补缺失值的常用方法是通过规则。然而,对于具有数千个缺失数据点的大型数据集,用户理解数据并制定有效的完成规则是费力且耗时的。因此,需要向用户显示对填写缺失字段影响最大的数据子集。此外,这些子集应该为用户提供足够的信息来进行更新。从大型数据集中选择最大化填充缺失数据的概率的子集在计算上是昂贵的。为了解决这些挑战,我们提出了ICARUS,它使用一个启发式算法,以矩阵的形式向用户显示数据库的小子集。这允许用户通过基于他们对矩阵的直接编辑应用建议的规则来迭代地填充数据。建议的规则通过使用数据库模式来推断层次结构,将用户的输入放大到多个缺失的字段。模拟结果显示,ICARUS在三个数据集上的平均性能比基线系统提高了50%。此外,面对面的用户研究表明,天真的用户可以在一个小时内填写68%的缺失数据,而手动规则规范需要几周的时间。
An important step in data preparation involves dealing with incomplete datasets. In some cases, the missing values are unreported because they are characteristics of the domain and are known by practitioners. Due to this nature of the missing values, imputation and inference methods do not work and input from domain experts is required. A common method for experts to fill missing values is through rules. However, for large datasets with thousands of missing data points, it is laborious and time consuming for a user to make sense of the data and formulate effective completion rules. Thus, users need to be shown subsets of the data that will have the most impact in completing missing fields. Further, these subsets should provide the user with enough information to make an update. Choosing subsets that maximize the probability of filling in missing data from a large dataset is computationally expensive. To address these challenges, we present ICARUS, which uses a heuristic algorithm to show the user small subsets of the database in the form of a matrix. This allows the user to iteratively fill in data by applying suggested rules based on their direct edits to the matrix. The suggested rules amplify the users’ input to multiple missing fields by using the database schema to infer hierarchies. Simulations show ICARUS has an average improvement of 50% across three datasets over the baseline system. Further, in-person user studies demonstrate that naive users can fill in 68% of missing data within an hour, while manual rule specification spans weeks.