Multi-way set enumeration in weight tensors

Multi-way set enumeration in weight tensors
复制标题

DOI:
10.1007/s10994-010-5210-y
复制
发表时间:
2011-02-01
期刊:
影响因子:
7.5
通讯作者:
Schoelkopf, Bernhard
Schoelkopf, Bernhard
中科院分区:
计算机科学3区
文献类型:
--
作者:
Georgii, Elisabeth;Tsuda, Koji;Schoelkopf, Bernhard

文献摘要

被引文献

相似文献

n-ary关系的分析在许多不同的领域受到关注,例如生物学,Web挖掘和社会研究。在基本设置中,有n组实例,每个观察关联n个实例,每组一个。探索这些n路数据的一种常见方法是搜索n集模式,即项集的n路等价物。更确切地说,n-集合模式由n个实例集合的特定子集组成,使得在数据中观察到对应实例之间的所有可能关联。相比之下,传统的项集挖掘方法只考虑双向数据,即项目与事务。n-set模式提供了更高级别的数据视图,揭示了实例组之间的关联关系。在这里,我们将这种方法概括为两个方面。首先,我们在一定程度上容忍缺失的观测值,这意味着我们也对n-集感兴趣,其中大多数(尽管不是全部)可能的关联都记录在数据中。其次,我们考虑关联权重。事实上,我们提出了一种方法来枚举所有的n-集,满足一个最小阈值的平均关联权重。从技术上讲,我们使用反向搜索策略来解决枚举任务,这允许有效地修剪搜索空间。此外,我们的算法提供了一个排名的解决方案,并可以考虑进一步的限制。我们展示了来自不同领域的人工和真实世界数据集上的实验结果。
The analysis of n-ary relations receives attention in many different fields, for instance biology, web mining, and social studies. In the basic setting, there are n sets of instances, and each observation associates n instances, one from each set. A common approach to explore these n-way data is the search for n-set patterns, the n-way equivalent of itemsets. More precisely, an n-set pattern consists of specific subsets of the n instance sets such that all possible associations between the corresponding instances are observed in the data. In contrast, traditional itemset mining approaches consider only two-way data, namely items versus transactions. The n-set patterns provide a higher-level view of the data, revealing associative relationships between groups of instances. Here, we generalize this approach in two respects. First, we tolerate missing observations to a certain degree, that means we are also interested in n-sets where most (although not all) of the possible associations have been recorded in the data. Second, we take association weights into account. In fact, we propose a method to enumerate all n-sets that satisfy a minimum threshold with respect to the average association weight. Technically, we solve the enumeration task using a reverse search strategy, which allows for effective pruning of the search space. In addition, our algorithm provides a ranking of the solutions and can consider further constraints. We show experimental results on artificial and real-world datasets from different domains.