A simple relevancy-ranking strategy for an interface to Boolean OPACs

A simple relevancy-ranking strategy for an interface to Boolean OPACs
复制标题

用于布尔 OPAC 接口的简单相关性排名策略

DOI:
--
复制
发表时间:
2004
期刊:
影响因子:
1.9
通讯作者:
K. Wan
K. Wan
中科院分区:
管理学4区
文献类型:
--
作者:
Christopher S. G. Khoo;K. Wan

文献摘要

被引文献

相似文献

为布尔在线公共访问目录(OPAC)的自然语言界面制定了一个相关性排名算法,并与作者正在开发的基于知识的搜索界面E-Referencer中目前使用的算法进行了比较。该算法使用了七个著名的排名标准:匹配广度,部分权重,查询词的接近度,变体词形式(词干),文档频率,术语频率和文档长度。该算法将自然语言查询转换为一系列越来越广泛的布尔搜索语句。在一个有10名受试者的小型实验中,该算法获得了良好的结果,平均整体精度为0.42,平均平均精度为0.62,与E-Referencer相比,精度提高了27%,平均精度提高了41%。分析了算法中每一步的有效性,并提出了改进建议。
A relevancy‐ranking algorithm for a natural language interface to Boolean online public access catalogs (OPACs) was formulated and compared with that currently used in a knowledge‐based search interface called the E‐Referencer, being developed by the authors. The algorithm makes use of seven well‐known ranking criteria: breadth of match, section weighting, proximity of query words, variant word forms (stemming), document frequency, term frequency and document length. The algorithm converts a natural language query into a series of increasingly broader Boolean search statements. In a small experiment with ten subjects in which the algorithm was simulated by hand, the algorithm obtained good results with a mean overall precision of 0.42 and mean average precision of 0.62, representing a 27 percent improvement in precision and 41 percent improvement in average precision compared to the E‐Referencer. The usefulness of each step in the algorithm was analyzed and suggestions are made for improving the algorithm.