Hybrid Context-Aware Word Sense Disambiguation in Topic Modeling based Document Representation

Hybrid Context-Aware Word Sense Disambiguation in Topic Modeling based Document Representation
复制标题

DOI:
10.1109/icdm50108.2020.00042
复制
发表时间:
2020-11
期刊:
2020 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Wenbo Li;Einoshin Suzuki
Wenbo Li;Einoshin Suzuki
中科院分区:
其他
文献类型:
--
作者:
Wenbo Li;Einoshin Suzuki

文献摘要

相似文献

我们提出了一种基于混合上下文的主题模型,用于文档表示中的词义消歧。文档表示是各种基于文档的任务的重要组成部分,词义消歧是为了捕获表示中词义的区别。传统方法主要依靠知识库进行数据充实;然而,单词的语义划分可能因不同的特定领域数据集而异。我们的目标是发现每个输入数据集更具体的单词语义差异,并在不丰富数据的情况下处理消歧问题。这种消歧的挑战是(1)为每个多义词划分不同的含义,同时(2)保留同义词之间的差异。大多数现有模型要么基于单独的上下文集群,要么集成辅助模块来指定词义。它们很难同时实现(1)和(2),因为一个词的不同含义被假定为独立的并且它们的内在关系被忽略。为了解决这个问题,我们通过词义出现的上下文和其他出现的上下文来估计词义。此外,我们引入了“Bag-of-Senses”(BoS)假设:文档是词义的多重集,并且生成词义而不是单词。我们对三个标准数据集的实验表明,我们的建议在词义估计、主题建模和文档分类的准确性方面优于其他最先进的方法。
We propose a hybrid context based topic model for word sense disambiguation in document representation. Document representation is an essential part of various document based tasks, and word sense disambiguation is to capture the distinctions of word senses in the representation. Traditional methods mainly rely on knowledge libraries for data enrichment; however, semantics division for a word may vary from different domain-specific datasets. We aim to discover more particular word semantic differences for each input dataset and handle the disambiguation problem without data enrichment. The challenge for this disambiguation is to (1) divide various senses for each polysemous word while (2) preserve the differences between synonyms. Most of the existing models are either based on separate context clusters or integrating an auxiliary module to specify word senses. They can hardly achieve both (1) and (2) since different senses of a word are assumed to be independent and their intrinsic relationships are ignored. To solve this problem, we estimate a word sense by both the context in which it occurs and the contexts of its other occurrences. Besides, we introduce the “Bag-of-Senses” (BoS) assumption: a document is a multiset of word senses, and the senses are generated instead of the words. Our experiments on three standard datasets show that our proposal outperforms other state-of-the-art methods in terms of accuracy of word sense estimation, topic modeling, and document classification.