An empirical study on retrieval models for different document genres: patents and newspaper articles

An empirical study on retrieval models for different document genres: patents and newspaper articles
复制标题

DOI:
10.1145/860435.860482
复制
发表时间:
2003-07
期刊:
Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval
影响因子:
--
通讯作者:
Makoto Iwayama;Atsushi Fujii;N. Kando;Yuzo Marukawa
Makoto Iwayama;Atsushi Fujii;N. Kando;Yuzo Marukawa
中科院分区:
其他
文献类型:
--
作者:
Makoto Iwayama;Atsushi Fujii;N. Kando;Yuzo Marukawa

文献摘要

被引文献

相似文献

自 20 世纪 90 年代以来,大量测试数据集在信息检索中的利用迅速增长,人们进行了大量的比较实验来探索各种检索模型的有效性。然而,大多数馆藏旨在检索报纸文章和技术摘要。在本文中,我们描述了制作专利检索测试集(NTCIR-3 专利检索集)的过程,其中包括两年的日本专利申请和由专业专利检索员制作的 31 个主题。我们还报告了使用该集合重新检验专利检索背景下现有检索模型的有效性所获得的实验结果。现有检索模型之间的相对优势并没有因文档类型(即专利和报纸文章)而显着差异。还讨论了与专利检索相关的问题。
Reflecting the rapid growth in the utilization of large test collections for information retrieval since the 1990s, extensive comparative experiments have been performed to explore the effectiveness of various retrieval models. However, most collections were intended for retrieving newspaper articles and technical abstracts. In this paper, we describe the process of producing a test collection for patent retrieval, the NTCIR-3 Patent Retrieval Collection, which includes two years of Japanese patent applications and 31 topics produced by professional patent searchers. We also report experimental results obtained by using this collection to re-examine the effectiveness of existing retrieval models in the context of patent retrieval. The relative superiority among existing retrieval models did not significantly differ depending on the document genre, that is, patents and newspaper articles. Issues related to patent retrieval are also discussed.