BATS: A Spectral Biclustering Approach to Single Document Topic Modeling and Segmentation

BATS: A Spectral Biclustering Approach to Single Document Topic Modeling and Segmentation
复制标题

DOI:
10.1145/3468268
复制
发表时间:
2021-12-01
影响因子:
5
通讯作者:
Li, Yanhua
Li, Yanhua
中科院分区:
计算机科学3区
文献类型:
--
作者:
Wu, Qiong;Hare, Adam;Li, Yanhua

文献摘要

被引文献

相似文献

现有的主题建模和文本分割方法通常需要大型数据集进行训练,当只有少量文本可用时,限制了它们的能力。在这项工作中,我们重新审视的“主题识别”和“文本分割”的稀疏文档学习的相互关联的问题,当有一个新的感兴趣的文本。在开发处理单个文档的方法时,我们面临两个主要挑战。首先是稀疏信息:只能访问一个文档,我们无法训练传统的主题模型或深度学习算法。第二个是显著噪声:任何单个文档中的相当大一部分单词只会产生噪声,而无助于识别主题或片段。为了解决这些问题,我们设计了一个无监督的,计算效率高的方法,称为双聚类方法主题建模和分割(BATS)。BATS利用三个关键思想来同时识别主题和分段文本:(i)使用词序信息来降低样本复杂性的新机制,(ii)基于统计学的基于图的双聚类技术,用于识别单词和句子的潜在结构,以及(iii)有效的分类方法的集合,用于去除噪声单词并奖励重要单词以进一步提高性能。六个数据集上的实验表明,我们的方法优于几个国家的最先进的基线时,考虑主题的一致性,主题的多样性,分割和运行时比较指标。
Existing topic modeling and text segmentation methodologies generally require large datasets for training, limiting their capabilities when only small collections of text are available. In this work, we reexamine the inter-related problems of "topic identification" and "text segmentation" for sparse document learning, when there is a single new text of interest. In developing a methodology to handle single documents, we face two major challenges. First is sparse information: with access to only one document, we cannot train traditional topic models or deep learning algorithms. Second is significant noise: a considerable portion of words in any single document will produce only noise and not help discern topics or segments. To tackle these issues, we design an unsupervised, computationally efficient methodology called Biclustering Approach to Topic modeling and Segmentation (BATS). BATS leverages three key ideas to simultaneously identify topics and segment text: (i) a new mechanism that uses word order information to reduce sample complexity, (ii) a statistically sound graph-based biclustering technique that identifies latent structures of words and sentences, and (iii) a collection of effective heuristics that remove noise words and award important words to further improve performance. Experiments on six datasets show that our approach outperforms several state-of-the-art baselines when considering topic coherence, topic diversity, segmentation, and runtime comparison metrics.