Fast Extraction of Semantic Features from a Latent Semantic Indexed Text Corpus

Fast Extraction of Semantic Features from a Latent Semantic Indexed Text Corpus
复制标题

从潜在语义索引文本语料库中快速提取语义特征

DOI:
10.1023/a:1013801028884
复制
发表时间:
2002
影响因子:
3.1
通讯作者:
M. Girolami
M. Girolami
中科院分区:
计算机科学4区
文献类型:
--
作者:
A. Kabán;M. Girolami

文献摘要

被引文献

相似文献

本文提出了一种基于投影的对称因子分解方法,用于从存储在潜在语义空间中的文本文档集合中提取语义特征。初步的实验结果表明,这产生了一个可比的表示提供了一种新的概率方法,重新考虑整个索引问题的文本文档,并直接在原来的高维向量空间表示的文本。所采用的投影指数在这里是从问题的先验约束导出的。这种方法的主要优点是计算效率,并通过利用潜在语义索引作为预处理阶段。在各种设置的20个新闻组文本语料库的子集上的模拟结果。
This paper proposes a projection-based symmetrical factorisation method for extracting semantic features from collections of text documents stored in a Latent Semantic space. Preliminary experimental results demonstrate this yields a comparable representation to that provided by a novel probabilistic approach which reconsiders the entire indexing problem of text documents and works directly in the original high dimensional vector-space representation of text. The employed projection index is derived here from thea prioriconstraints on the problem. The principal advantage of this approach is computational efficiency and is obtained by the exploitation of the Latent Semantic Indexing as a preprocessing stage. Simulation results on subsets of the 20-Newsgroups text corpus in various settings are provided.