PROBABILISTIC APPROACH TO AUTOMATIC KEYWORD INDEXING .1. DISTRIBUTION OF SPECIALTY WORDS IN A TECHNICAL LITERATURE

PROBABILISTIC APPROACH TO AUTOMATIC KEYWORD INDEXING .1. DISTRIBUTION OF SPECIALTY WORDS IN A TECHNICAL LITERATURE
复制标题

DOI:
10.1002/asi.4630260402
复制
发表时间:
1975-01-01
期刊:
JOURNAL OF THE AMERICAN SOCIETY FOR INFORMATION SCIENCE
影响因子:
--
通讯作者:
HARTER, SP
HARTER, SP
中科院分区:
其他
文献类型:
--
作者:
HARTER, SP

文献摘要

被引文献

相似文献

本研究研究的问题是开发一套正式的统计规则,用于识别文档的关键字-可能作为该文档索引术语有用的单词。这项研究是由许多作者观察到的,即非专业词汇,对索引目的没有什么价值的词汇,倾向于随机分布在文档集合中。相比之下,专业词汇的分布就不那么均匀了。在研究的第一部分中,我们详细研究了两个泊松分布的混合模型作为专门词分布的模型,并推导了用经验频率统计量表示该模型的三个参数的公式。模型的拟合在实验文件集上进行了测试,并发现可以接受研究的目的。一种旨在识别专业词汇的措施,与2‐泊松模型一致,提出并评估。
The problem studied in this research is that of developing a set of formal statistical rules for the purpose of identifying the keywords of a document‐words likely to be useful as index terms for that document. The research was prompted by the observation, made by a number of writers, that non‐specialty words, words which possess little value for indexing purposes, tend to be distributed at random in a collection of documents. In contrast, specialty words are not so distributed.In Part I of the study, a mixture of two Poisson distributions is examined in detail as a model of specialty word distribution, and formulas expressing the three parameters of the model in terms of empirical frequency statistics are derived. The fit of the model is tested on an experimental document collection and found to be acceptable for the purposes of the study. A measure intended to identify specialty words, consistent with the 2‐Poisson model, is proposed and evaluated.