Never Abandon Minorities: Exhaustive Extraction of Bursty Phrases on Microblogs Using Set Cover Problem

Never Abandon Minorities: Exhaustive Extraction of Bursty Phrases on Microblogs Using Set Cover Problem
复制标题

DOI:
10.18653/v1/d17-1251
复制
发表时间:
2017-09
期刊:
2017 IEEE/ACM Second International Conference on Internet-of-Things Design and Implementation (IoTDI)
影响因子:
--
通讯作者:
Masumi Shirakawa;T. Hara;T. Maekawa
Masumi Shirakawa;T. Hara;T. Maekawa
中科院分区:
其他
文献类型:
--
作者:
Masumi Shirakawa;T. Hara;T. Maekawa

文献摘要

相似文献

我们提出了一种独立于语言的数据驱动的方法来详尽地提取任意形式的突发短语(例如,短语而不是简单的名词短语)。突发(即,短语的出现的快速增加)导致包括不完整的N元语法的重叠N元语法的突发。换句话说,突发不完整的N元语法不可避免地与突发短语重叠。因此,所提出的方法执行突发短语的提取作为集合覆盖问题,其中所有突发N-gram被突发短语的最小集合覆盖。使用日本Twitter数据的实验结果表明,所提出的方法优于基于词,名词短语为基础的,基于分割的方法在准确性和覆盖率。
We propose a language-independent data-driven method to exhaustively extract bursty phrases of arbitrary forms (e.g., phrases other than simple noun phrases) from microblogs. The burst (i.e., the rapid increase of the occurrence) of a phrase causes the burst of overlapping N-grams including incomplete ones. In other words, bursty incomplete N-grams inevitably overlap bursty phrases. Thus, the proposed method performs the extraction of bursty phrases as the set cover problem in which all bursty N-grams are covered by a minimum set of bursty phrases. Experimental results using Japanese Twitter data showed that the proposed method outperformed word-based, noun phrase-based, and segmentation-based methods both in terms of accuracy and coverage.