Exploratory Analysis of Graph Data by Leveraging Domain Knowledge

Exploratory Analysis of Graph Data by Leveraging Domain Knowledge
复制标题

DOI:
10.1109/icdm.2017.28
复制
发表时间:
2017-11
期刊:
2017 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Di Jin;Danai Koutra
Di Jin;Danai Koutra
中科院分区:
其他
文献类型:
--
作者:
Di Jin;Danai Koutra

文献摘要

被引文献

相似文献

鉴于每天生成的数据量激增,图挖掘任务变得越来越具有挑战性,导致对摘要技术的巨大需求。特征选择是一种代表性的方法,它通过选择与特定任务(如分类,预测和异常检测)相关的特征来简化数据集。虽然它可以被看作是一种根据一些特征来总结图的方式,但它并没有很好地定义用于探索性分析,并且它联合而不是有条件地对一组观察结果进行操作(即,从许多图中选择特征相对于以其它图为条件的输入图的选择)。在这项工作中,我们介绍了EAGLE(探索性分析图与域knowLEdge),一种新的方法,创建可解释的,基于特征的,和特定于域的图形摘要在一个完全自动的方式。也就是说,不同域中的相同图-例如,社会科学和神经科学将通过不同的EAGLE摘要进行描述,这些摘要自动利用领域知识和期望。我们提出了一个优化公式,旨在找到一个可解释的摘要与最具代表性的功能的输入图,使它是:多样的,简洁的,特定领域的,和高效的。在合成数据集和真实数据集上进行的大量实验表明,EAGLE的有效性和效率以及其优于现有方法的优势。我们还展示了我们的方法可以应用于各种图挖掘任务,如分类和探索性分析。
Given the soaring amount of data being generated daily, graph mining tasks are becoming increasingly challenging, leading to tremendous demand for summarization techniques. Feature selection is a representative approach that simplifies a dataset by choosing features that are relevant to a specific task, such as classification, prediction, and anomaly detection. Although it can be viewed as a way to summarize a graph in terms of a few features, it is not well-defined for exploratory analysis, and it operates on a set of observations jointly rather than conditionally (i.e., feature selection from many graphs vs. selection for an input graph conditioned on other graphs). In this work, we introduce EAGLE (Exploratory Analysis of Graphs with domain knowLEdge), a novel method that creates interpretable, feature-based, and domain-specific graph summaries in a fully automatic way. That is, the same graph in different domains–e.g., social science and neuroscience–will be described via different EAGLE summaries, which automatically leverage the domain knowledge and expectations. We propose an optimization formulation that seeks to find an interpretable summary with the most representative features for the input graph so that it is: diverse, concise, domain-specific, and efficient. Extensive experiments on synthetic and real-world datasets with up to ~1M edges and ~400 features demonstrate the effectiveness and efficiency of EAGLE and its benefits over existing methods. We also show how our method can be applied to various graph mining tasks, such as classification and exploratory analysis.