Making interval-based clustering rank-aware

Making interval-based clustering rank-aware
复制标题

使基于间隔的聚类具有排名意识

DOI:
--
复制
发表时间:
2011
期刊:
EDBT/ICDT Workshops
影响因子:
--
通讯作者:
Tova Milo
Tova Milo
中科院分区:
--
文献类型:
--
作者:
Julia Stoyanovich;S. Amer;Tova Milo

文献摘要

被引文献

相似文献

在在线应用程序(例如在线约会)中,用户经常查询和对大量结构化项目进行排名。最佳结果往往是均匀的,这阻碍了数据探索。例如,一个正在寻找20至40岁伙伴的约会网站用户,并通过收入从更高到更低的收入进行比赛,将在30年代后期拥有大量的比赛,他们拥有MBA学位和工作在金融业,在看到不同年龄段和各行各业的任何比赛之前。在排名列表中呈现结果的一种替代方法是在结果空间中找到簇,并通过与等级相关的属性组合确定。这样的集群可以用MBA的35到40个匹配,在软件行业工作的25至30岁之间的匹配项,从而可以进行排名结果的数据探索。 我们指的是找到这些簇的问题,例如基于等级的间隔聚类,并认为它不是通过标准聚类算法来解决的。我们正式定义了该问题,并解决问题,提出了一种新型的局部衡量标准,以及一个适合本应用方案的聚类质量度量。这些成分可以由多种聚类算法使用,我们提出了BARAC,这是一种特定的子空间聚类算法,可以在具有异质属性的域中基于等级的间隔聚类。我们通过大规模的用户研究来验证方法的有效性,并对效率进行广泛的实验评估,表明我们的方法在大规模上是实用的。我们的评估是在Yahoo!的大型数据集上进行的。交友,领先的在线约会网站以及Yahoo!的餐厅数据当地的。
In online applications, such as online dating, users often query and rank large collections of structured items. Top results tend to be homogeneous, which hinders data exploration. For example, a dating website user who is looking for a partner between 20 and 40 years old, and who sorts the matches by income from higher to lower, will see a large number of matches in their late 30s who hold an MBA degree and work in the financial industry, before seeing any matches in different age groups and walks of life. An alternative to presenting results in a ranked list is to find clusters in the result space, identified by a combination of attributes that correlate with rank. Such clusters may describe matches between 35 and 40 with an MBA, matches between 25 and 30 who work in the software industry, etc., allowing for data exploration of ranked results. We refer to the problem of finding such clusters as rank-aware interval-based clustering and argue that it is not addressed by standard clustering algorithms. We formally define the problem and, to solve it, propose a novel measure of locality, together with a family of clustering quality measures appropriate for this application scenario. These ingredients may be used by a variety of clustering algorithms, and we present BARAC, a particular subspace-clustering algorithm that enables rank-aware interval-based clustering in domains with heterogeneous attributes. We validate the effectiveness of our approach with a large-scale user study, and perform an extensive experimental evaluation of efficiency, demonstrating that our methods are practical on the large scale. Our evaluation is performed on large datasets from Yahoo! Personals, a leading online dating site, and on restaurant data from Yahoo! Local.