THE AUTOMATIC IDENTIFICATION OF STOP WORDS

THE AUTOMATIC IDENTIFICATION OF STOP WORDS
复制标题

DOI:
10.1177/016555159201800106
复制
发表时间:
1992-01-01
影响因子:
2.4
通讯作者:
SIROTKIN, K
SIROTKIN, K
中科院分区:
计算机科学3区
文献类型:
--
作者:
WILBUR, WJ;SIROTKIN, K

文献摘要

被引文献

相似文献

停用词可被标识为在与查询无关的那些文档中出现的可能性与在与查询相关的那些文档中出现的可能性相同的词。在这篇文章中,我们展示了相关性的概念如何被相似度高的条件所取代。因此,可以通过自动统计测试来识别集合中的停用词。我们描述了统计测试的性质,因为它是通过基于文档-文档相似度的余弦系数的矢量检索方法实现的。例如,该技术随后被应用于生物技术领域中的大型MEDLINE(R)子集。该数据库的初始处理包括310个单词的常用非内容术语的停顿列表。然后应用我们的技术,剩余术语的75%被识别为停用词。我们比较了删除这些停用词和不删除这些停用词的检索,并发现在响应随机查询文档而检索到的前20个文档中,对于这两种方法来说,其中17个文档的平均值相同。我们还检查了差异并得出结论,在用户更喜欢一种方法而不是另一种方法的情况下,具有精简术语集的新方法大约有四分之三受到青睐。
A stop word may be identified as a word that has the same likelihood of occurring in those documents not relevant to a query as in those documents relevant to the query. In this paper we show how the concept of relevance may be replaced by the condition of being highly rated by a similarity measure. Thus it becomes possible to identify the stop words in a collection by automated statistical testing. We describe the nature of the statistical test as it is realized with a vector retrieval methodology based on the cosine coefficient of document-document similarity. As an example, this technique is then applied to a large MEDLINE(R) subset in the area of biotechnology. The initial processing of this database involves a 310 word stop list of common non-content terms. Our technique is then applied and 75% of the remaining terms are identified as stop words. We compare retrieval with and without the removal of these stop words and find that of the top twenty documents retrieved in response to a random query document, seventeen of these are the same on the average for the two methods. We also examine the differences and conclude that where the user prefers one method over the other, the new method with the reduced term set is favored about three times out of four.