MetaPAD: Meta Pattern Discovery from Massive Text Corpora

MetaPAD: Meta Pattern Discovery from Massive Text Corpora
复制标题

DOI:
10.1145/3097983.3098105
复制
发表时间:
2017-03
期刊:
Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Meng Jiang;Jingbo Shang;Taylor Cassidy;Xiang Ren;Lance M. Kaplan;T. Hanratty;Jiawei Han
Meng Jiang;Jingbo Shang;Taylor Cassidy;Xiang Ren;Lance M. Kaplan;T. Hanratty;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Meng Jiang;Jingbo Shang;Taylor Cassidy;Xiang Ren;Lance M. Kaplan;T. Hanratty;Jiawei Han

文献摘要

被引文献

相似文献

在新闻、推特、论文等文本语料库中挖掘文本模式一直是文本挖掘和自然语言处理研究中的一个活跃主题。以前的研究采用依赖分析为基础的模式发现方法。然而,分析结果失去了丰富的上下文周围的实体的模式,这一过程是昂贵的大规模语料库。在本研究中,我们提出了一种新的类型化的文本模式结构,称为Meta模式,它是在一定的背景下扩展到一个频繁的,信息丰富的,和精确的子序列模式。本文提出了一个高效的元模式发现框架MetaPAD,该框架通过三种技术从海量语料库中发现Meta模式:(1)提出了一种上下文感知的模式分割方法,利用学习的模式质量评估函数仔细地确定模式的边界,避免了昂贵的依赖分析,生成高质量的模式;(2)从同义元模式的类型、语境和提取等方面对同义Meta模式进行识别和分类;以及(3)检查由每组模式提取的实例中的实体的类型分布,并寻找合适的类型级别以使发现的模式精确。实验表明,我们提出的框架发现高质量的类型化文本模式有效地从不同类型的海量语料库,促进信息提取。
Mining textual patterns in news, tweets, papers, and many other kinds of text corpora has been an active theme in text mining and NLP research. Previous studies adopt a dependency parsing-based pattern discovery approach. However, the parsing results lose rich context around entities in the patterns, and the process is costly for a corpus of large scale. In this study, we propose a novel typed textual pattern structure, called meta pattern, which is extended to a frequent, informative, and precise subsequence pattern in certain context. We propose an efficient framework, called MetaPAD, which discovers meta patterns from massive corpora with three techniques: (1) it develops a context-aware segmentation method to carefully determine the boundaries of patterns with a learnt pattern quality assessment function, which avoids costly dependency parsing and generates high-quality patterns; (2) it identifies and groups synonymous meta patterns from multiple facets---their types, contexts, and extractions; and (3) it examines type distributions of entities in the instances extracted by each group of patterns, and looks for appropriate type levels to make discovered patterns precise. Experiments demonstrate that our proposed framework discovers high-quality typed textual patterns efficiently from different genres of massive corpora and facilitates information extraction.