Effects of Stop Words Elimination for Arabic Information Retrieval: A Comparative Study

Effects of Stop Words Elimination for Arabic Information Retrieval: A Comparative Study
复制标题

停用词消除对阿拉伯语信息检索的影响:比较研究

DOI:
--
复制
发表时间:
2017
期刊:
ArXiv
影响因子:
--
通讯作者:
I. A. El
I. A. El
中科院分区:
--
文献类型:
--
作者:
I. A. El

文献摘要

被引文献

相似文献

本研究探讨三种停用词表-通用停用词表、语料库停用词表、组合停用词表-在阿拉伯文信息检索中的效能。三个流行的加权方案进行了检查:逆文档频率权重,概率加权,统计语言建模。其思想是将联合收割机的统计方法和语言方法结合起来,以达到最佳的性能,并比较它们在检索上的效果。最不发达国家(语言数据联盟)阿拉伯语新闻网数据集与狐猴工具包一起使用。Okapi检索系统中使用的最佳匹配加权方案在研究中使用的三种加权算法中具有最佳的整体性能,stoplists提高了检索效率,特别是当与BM 25权重一起使用时。一般非索引字列表的整体表现优于其他两个列表。
The effectiveness of three stop words lists for Arabic Information Retrieval---General Stoplist, Corpus- Based Stoplist, Combined Stoplist ---were investigated in this study. Three popular weighting schemes were examined: the inverse document frequency weight, probabilistic weighting, and statistical language modelling. The Idea is to combine the statistical approaches with linguistic approaches to reach an optimal performance, and compare their effect on retrieval. The LDC (Linguistic Data Consortium) Arabic Newswire data set was used with the Lemur Toolkit. The Best Match weighting scheme used in the Okapi retrieval system had the best overall performance of the three weighting algorithms used in the study, stoplists improved retrieval effectiveness especially when used with the BM25 weight. The overall performance of a general stoplist was better than the other two lists.