TaxoGen: Unsupervised Topic Taxonomy Construction by Adaptive Term Embedding and Clustering

TaxoGen: Unsupervised Topic Taxonomy Construction by Adaptive Term Embedding and Clustering
复制标题

DOI:
10.1145/3219819.3220064
复制
发表时间:
2018-07
期刊:
Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子:
--
通讯作者:
Chao Zhang;Fangbo Tao;Xiusi Chen;Jiaming Shen;Meng Jiang;Brian M. Sadler;M. Vanni;Jiawei Han
Chao Zhang;Fangbo Tao;Xiusi Chen;Jiaming Shen;Meng Jiang;Brian M. Sadler;M. Vanni;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Chao Zhang;Fangbo Tao;Xiusi Chen;Jiaming Shen;Meng Jiang;Brian M. Sadler;M. Vanni;Jiawei Han

文献摘要

被引文献

相似文献

分类结构的构建不仅是文本语义分析的基础工作,也是信息过滤、推荐、Web搜索等应用的重要步骤。现有的基于模式的方法提取上下义词对,然后将这些对组织成一个分类。然而,由于将每个术语视为一个独立的概念节点,它们忽略了术语之间的主题接近性和语义相关性。在本文中,我们提出了一种构建主题分类的方法,其中每个节点表示一个概念主题,并定义为一个语义一致的概念术语的集群。我们的方法,TaxoGen,使用术语嵌入和层次聚类,以递归的方式构建一个主题分类。为了确保递归过程的质量,它包括:(1)自适应球形聚类模块,用于在将粗主题划分为细粒度主题时将术语分配到适当的级别;(2)局部嵌入模块,用于学习在分类的不同级别保持强区分力的术语嵌入。我们在两个真实的数据集上的实验证明了与基线方法相比TaxoGen的有效性。
Taxonomy construction is not only a fundamental task for semantic analysis of text corpora, but also an important step for applications such as information filtering, recommendation, and Web search. Existing pattern-based methods extract hypernym-hyponym term pairs and then organize these pairs into a taxonomy. However, by considering each term as an independent concept node, they overlook the topical proximity and the semantic correlations among terms. In this paper, we propose a method for constructing topic taxonomies, wherein every node represents a conceptual topic and is defined as a cluster of semantically coherent concept terms. Our method, TaxoGen, uses term embeddings and hierarchical clustering to construct a topic taxonomy in a recursive fashion. To ensure the quality of the recursive process, it consists of: (1) an adaptive spherical clustering module for allocating terms to proper levels when splitting a coarse topic into fine-grained ones; (2) a local embedding module for learning term embeddings that maintain strong discriminative power at different levels of the taxonomy. Our experiments on two real datasets demonstrate the effectiveness of TaxoGen compared with baseline methods.