Complete mining of frequent patterns from graphs: Mining graph data

Complete mining of frequent patterns from graphs: Mining graph data
复制标题

DOI:
10.1023/a:1021726221443
复制
发表时间:
2003-03-01
期刊:
影响因子:
7.5
通讯作者:
Motoda, H
Motoda, H
中科院分区:
计算机科学3区
文献类型:
--
作者:
Inokuchi, A;Washio, T;Motoda, H

文献摘要

被引文献

相似文献

篮子分析是一种标准的数据挖掘方法,它从数据库中提取频繁项集。然而,它的挖掘能力仅限于由项目组成的交易数据。在现实中,有许多应用程序中的数据是以更结构化的方式描述的,例如。G.化合物和网页浏览历史。在机器学习领域,有一些方法可以从图结构数据中发现特征模式。然而,几乎所有这些都不适合需要完整搜索数据中所有频繁子图模式的应用。本文提出了一种新的原理和算法,用于提取图结构数据中频繁出现的特征模式。我们的算法可以从有向和无向图结构数据中导出所有频繁诱导子图,这些数据具有带有标记或未标记节点和链接的循环(包括自循环)。它的性能进行了评估,通过应用到Web浏览模式分析和化学致癌分析。
Basket Analysis, which is a standard method for data mining, derives frequent itemsets from database. However, its mining ability is limited to transaction data consisting of items. In reality, there are many applications where data are described in a more structural way, e. g. chemical compounds and Web browsing history. There are a few approaches that can discover characteristic patterns from graph-structured data in the field of machine learning. However, almost all of them are not suitable for such applications that require a complete search for all frequent subgraph patterns in the data. In this paper, we propose a novel principle and its algorithm that derive the characteristic patterns which frequently appear in graph-structured data. Our algorithm can derive all frequent induced subgraphs from both directed and undirected graph structured data having loops (including self-loops) with labeled or unlabeled nodes and links. Its performance is evaluated through the applications to Web browsing pattern analysis and chemical carcinogenesis analysis.