Regret-minimizing representative databases

Regret-minimizing representative databases
复制标题

DOI:
10.14778/1920841.1920980
复制
发表时间:
2010-09
影响因子:
2.5
通讯作者:
Danupon Nanongkai;Atish Das Sarma;Ashwin Lall;R. Lipton;Jun Xu
Danupon Nanongkai;Atish Das Sarma;Ashwin Lall;R. Lipton;Jun Xu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Danupon Nanongkai;Atish Das Sarma;Ashwin Lall;R. Lipton;Jun Xu

文献摘要

被引文献

相似文献

我们提出了k-代表遗憾最小化查询(k-遗憾)作为一个操作,以支持多标准决策。与top-k类似,k-regret查询假设用户具有一些效用或评分函数;然而,它从不要求用户提供这些函数。像skyline一样,它根据用户的标准从一个潜在的大型数据库中过滤出一组感兴趣的点;然而,它不会通过输出太多的元组来干扰用户。特别地,对于任何数量k和任何类别的效用函数,k-后悔查询从数据库中输出k个元组,并试图最小化最大后悔率。这捕获了如果用户看到k个代表元组而不是整个数据库,她会有多失望。我们专注于一类线性效用函数,这是广泛适用的。这种方法的第一个挑战是,不清楚最大后悔率是否会很小,甚至有界。我们肯定地回答这个问题。理论上,我们证明了最大后悔率是有界的,并且这个界与数据库的大小无关。此外,我们在真实的和合成数据集上进行的大量实验表明,在实践中,最大后悔率相当小。另外,本文提出的算法在数据库规模下的运行时间是线性的,具有一定的实用性,实验表明,在skyline操作上运行时,算法的运行时间很小,可以集成到现有的数据库系统中。
We propose the k-representative regret minimization query (k-regret) as an operation to support multi-criteria decision making. Like top-k, the k-regret query assumes that users have some utility or scoring functions; however, it never asks the users to provide such functions. Like skyline, it filters out a set of interesting points from a potentially large database based on the users' criteria; however, it never overwhelms the users by outputting too many tuples. In particular, for any number k and any class of utility functions, the k-regret query outputs k tuples from the database and tries to minimize the maximum regret ratio. This captures how disappointed a user could be had she seen k representative tuples instead of the whole database. We focus on the class of linear utility functions, which is widely applicable. The first challenge of this approach is that it is not clear if the maximum regret ratio would be small, or even bounded. We answer this question affirmatively. Theoretically, we prove that the maximum regret ratio can be bounded and this bound is independent of the database size. Moreover, our extensive experiments on real and synthetic datasets suggest that in practice the maximum regret ratio is reasonably small. Additionally, algorithms developed in this paper are practical as they run in linear time in the size of the database and the experiments show that their running time is small when they run on top of the skyline operation which means that these algorithm could be integrated into current database systems.