Fast Extraction of Semantic Features from a Latent Semantic Indexed Text Corpus
Fast Extraction of Semantic Features from a Latent Semantic Indexed Text Corpus
复制标题
从潜在语义索引文本语料库中快速提取语义特征
DOI:
10.1023/a:1013801028884
复制
发表时间:
2002
影响因子:
3.1
通讯作者:
M. Girolami
中科院分区:
文献类型:
--
作者:
A. Kabán;M. Girolami
This paper proposes a projection-based symmetrical factorisation method for extracting semantic features from collections of text documents stored in a Latent Semantic space. Preliminary experimental results demonstrate this yields a comparable representation to that provided by a novel probabilistic approach which reconsiders the entire indexing problem of text documents and works directly in the original high dimensional vector-space representation of text. The employed projection index is derived here from thea prioriconstraints on the problem. The principal advantage of this approach is computational efficiency and is obtained by the exploitation of the Latent Semantic Indexing as a preprocessing stage. Simulation results on subsets of the 20-Newsgroups text corpus in various settings are provided.