Adaptive Concept Resolution for Document Representation and Its Applications in Text Mining

Adaptive Concept Resolution for Document Representation and Its Applications in Text Mining
复制标题

文档表示的自适应概念解析及其在文本挖掘中的应用

DOI:
10.1016/j.knosys.2014.10.003
复制
发表时间:
2015
影响因子:
8.8
通讯作者:
Shoaib Jameel
Shoaib Jameel
中科院分区:
计算机科学1区
文献类型:
--
作者:
Lidong Bing;Shan Jiang;Wai Lam;Yan Zhang;Shoaib Jameel

文献摘要

相似文献

在计算文档间的相似度时,同义词和多义词往往会带来一定的噪声。现有的基于本体的文档表示方法是静态的,使得用于表示文档的所选择的语义概念具有固定的分辨率。因此,它们不适应文档收集的特点和文本挖掘问题。我们提出了一个自适应概念解析(ACR)模型来克服这个问题。ACR可以从本体中学习概念边界,同时考虑特定文档集合的特性。然后,该边界为来自同一领域的文档提供定制的语义概念表示。ACR的另一个优点是,它适用于在训练文档集中给定组的分类任务和没有组信息可用的聚类任务。实验结果表明,ACR在几乎所有情况下都优于现有的静态方法。我们还提出了一种方法来集成维基百科的实体到一个专家编辑的本体,即WordNet,生成一个增强的本体命名为WordNet-Plus,其性能也检查下的ACR模型。由于高覆盖率,WordNet-Plus在分类中具有更多新鲜文档的数据集上的表现优于WordNet。
It is well-known that synonymous and polysemous terms often bring in some noise when we calculate the similarity between documents. Existing ontology-based document representation methods are static so that the selected semantic concepts for representing a document have a fixed resolution. Therefore, they are not adaptable to the characteristics of document collection and the text mining problem in hand. We propose an Adaptive Concept Resolution (ACR) model to overcome this problem. ACR can learn a concept border from an ontology taking into the consideration of the characteristics of the particular document collection. Then, this border provides a tailor-made semantic concept representation for a document coming from the same domain. Another advantage of ACR is that it is applicable in both classification task where the groups are given in the training document set and clustering task where no group information is available. The experimental results show that ACR outperforms an existing static method in almost all cases. We also present a method to integrate Wikipedia entities into an expert-edited ontology, namely WordNet, to generate an enhanced ontology named WordNet-Plus, and its performance is also examined under the ACR model. Due to the high coverage, WordNet-Plus can outperform WordNet on data sets having more fresh documents in classification.