Applying Data Mining to Pseudo-Relevance Feedback for High Performance Text Retrieval

Applying Data Mining to Pseudo-Relevance Feedback for High Performance Text Retrieval
复制标题

DOI:
10.1109/icdm.2006.22
复制
发表时间:
2006-12
期刊:
Sixth International Conference on Data Mining (ICDM'06)
影响因子:
--
通讯作者:
Xiangji Huang;Y. Huang;M. Wen;Aijun An;Y. Liu;Josiah Poon
Xiangji Huang;Y. Huang;M. Wen;Aijun An;Y. Liu;Josiah Poon
中科院分区:
其他
文献类型:
--
作者:
Xiangji Huang;Y. Huang;M. Wen;Aijun An;Y. Liu;Josiah Poon

文献摘要

被引文献

相似文献

在本文中,我们研究了数据挖掘的使用,特别是文本分类和联合训练技术,以基于从检索系统的盲反馈获得的一小部分标记的段落来识别更相关的段落。数据挖掘结果用于扩展查询术语并重新估计在概率加权函数中使用的一些参数。我们在TREC硬数据集上对基于数据挖掘的反馈方法进行了评估。结果表明,数据挖掘可以成功地应用于提高文本检索的性能。我们详细地报告了我们的实验结果。
In this paper, we investigate the use of data mining, in particular the text classification and co-training techniques, to identify more relevant passages based on a small set of labeled passages obtained from the blind feedback of a retrieval system. The data mining results are used to expand query terms and to re-estimate some of the parameters used in a probabilistic weighting function. We evaluate the data mining based feedback method on the TREC HARD data set. The results show that data mining can be successfully applied to improve the text retrieval performance. We report our experimental findings in detail.