Indexing by Latent Semantic Analysis

Indexing by Latent Semantic Analysis
复制标题

DOI:
10.1002/(sici)1097-4571(199009)41:6
复制
发表时间:
1990-09
期刊:
J. Am. Soc. Inf. Sci.
影响因子:
--
通讯作者:
S. Deerwester;S. Dumais;T. Landauer;G. Furnas;R. Harshman
S. Deerwester;S. Dumais;T. Landauer;G. Furnas;R. Harshman
中科院分区:
其他
文献类型:
--
作者:
S. Deerwester;S. Dumais;T. Landauer;G. Furnas;R. Harshman

文献摘要

被引文献

相似文献

描述了一种自动索引和检索的新方法。该方法是利用术语与文档关联中的隐式高阶结构(“语义结构”),以便根据查询中找到的术语改进相关文档的检测。使用的特定技术是奇异值分解,其中将文档矩阵的大术语分解为一组大约。 100 个正交因子,可以通过线性组合近似原始矩阵。文档由 ca 表示。 100 个因子权重的项向量。查询被表示为由术语的加权组合形成的伪文档向量,并且返回具有超阈值余弦值的文档。初步测试发现这种完全自动的检索方法很有前途。
A new method for automatic indexing and retrieval is described. The approach is to take advantage of implicit higher-order structure in the association of terms with documents (“semantic structure”) in order to improve the detection of relevant documents on the basis of terms found in queries. The particular technique used is singular-value decomposition, in which a large term by document matrix is decomposed into a set of ca. 100 orthogonal factors from which the original matrix can be approximated by linear combination. Documents are represented by ca. 100 item vectors of factor weights. Queries are represented as pseudo-document vectors formed from weighted combinations of terms, and documents with supra-threshold cosine values are returned. initial tests find this completely automatic method for retrieval to be promising.