VECTOR-SPACE MODEL FOR AUTOMATIC INDEXING

VECTOR-SPACE MODEL FOR AUTOMATIC INDEXING
复制标题

DOI:
10.1145/361219.361220
复制
发表时间:
1975-01-01
影响因子:
22.7
通讯作者:
YANG, CS
YANG, CS
中科院分区:
计算机科学3区
文献类型:
--
作者:
SALTON, G;WONG, A;YANG, CS

文献摘要

被引文献

相似文献

在文档检索或其他模式匹配环境中,存储的实体(文档)相互比较或与传入的模式(搜索请求)进行比较,似乎最好的索引(属性)空间是每个实体尽可能远离其他实体的空间;在这种情况下,索引系统的值可以表示为对象空间密度的函数;特别是,检索性能可能与空间密度成反比。使用基于空间密度计算的方法为文档集合选择最佳索引词汇表。给出了典型的评价结果,说明了该模型的有效性。
In a document retrieval, or other pattern matching environment where stored entities (documents) are compared with each other or with incoming patterns (search requests), it appears that the best indexing (property) space is one where each entity lies as far away from the others as possible; in these circumstances the value of an indexing system may be expressible as a function of the density of the object space; in particular, retrieval performance may correlate inversely with space density. An approach based on space density computations is used to choose an optimum indexing vocabulary for a collection of documents. Typical evaluation results are shown, demonstating the usefulness of the model.