NetTaxo: Automated Topic Taxonomy Construction from Text-Rich Network

NetTaxo: Automated Topic Taxonomy Construction from Text-Rich Network
复制标题

DOI:
10.1145/3366423.3380259
复制
发表时间:
2020-04
期刊:
Proceedings of The Web Conference 2020
影响因子:
--
通讯作者:
Jingbo Shang;Xinyang Zhang;Liyuan Liu;Sha Li;Jiawei Han
Jingbo Shang;Xinyang Zhang;Liyuan Liu;Sha Li;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Jingbo Shang;Xinyang Zhang;Liyuan Liu;Sha Li;Jiawei Han

文献摘要

被引文献

相似文献

主题分类的自动化构造可以使许多应用受益,包括web搜索、推荐和知识发现。自动分类法构建的主要优点之一是能够捕获特定于语料库的信息并适应不同的场景。为了更好地反映语料库的特点,我们考虑了文档的元数据,并将语料库视为一个文本丰富的网络。在本文中,我们提出了NetTaxo,一个新的自动主题分类构建框架,它超越了现有的范式,并允许文本数据与网络结构的合作。具体来说,我们从文本和网络作为上下文来学习术语嵌入。采用网络基元来捕获适当的网络上下文。我们进行了一个实例级的选择图案,进一步细化根据每个分类节点的粒度和语义的术语嵌入。然后应用聚类以获得分类节点下的子主题。在两个真实数据集上的大量实验证明了该方法的优越性,进一步验证了实例级模体选择的有效性和重要性。
The automated construction of topic taxonomies can benefit numerous applications, including web search, recommendation, and knowledge discovery. One of the major advantages of automatic taxonomy construction is the ability to capture corpus-specific information and adapt to different scenarios. To better reflect the characteristics of a corpus, we take the meta-data of documents into consideration and view the corpus as a text-rich network. In this paper, we propose NetTaxo, a novel automatic topic taxonomy construction framework, which goes beyond the existing paradigm and allows text data to collaborate with network structure. Specifically, we learn term embeddings from both text and network as contexts. Network motifs are adopted to capture appropriate network contexts. We conduct an instance-level selection for motifs, which further refines term embedding according to the granularity and semantics of each taxonomy node. Clustering is then applied to obtain sub-topics under a taxonomy node. Extensive experiments on two real-world datasets demonstrate the superiority of our method over the state-of-the-art, and further verify the effectiveness and importance of instance-level motif selection.