Training query filtering for semi-supervised learning to rank with pseudo labels

Training query filtering for semi-supervised learning to rank with pseudo labels
复制标题

训练查询过滤以进行半监督学习以使用伪标签进行排名

DOI:
10.1007/s11280-015-0363-z
复制
发表时间:
2016-09
影响因子:
3.7
通讯作者:
Luo Tiejian
Luo Tiejian
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zhang Xin;He Ben;Luo Tiejian

文献摘要

参考文献

相似文献

半监督学习是一种机器学习范式,可用于在只有有限或没有可用训练示例的情况下从未标记的数据创建伪标签,以学习排名模型。然而,低质量的伪标签会阻碍半监督学习在信息检索(IR)中的有效性,因此需要训练查询过滤来去除低质量的查询。在本文中,我们假设两个应用场景方面的可用性的人类标签。首先,对于没有任何标记数据的应用程序,提出了一种基于聚类的方法来选择高质量的训练查询。该方法根据高质量训练查询的相关文档具有高度一致性的经验观察来选择训练查询。其次,对于有限的标记数据的应用程序,提出了一种基于分类的方法。该方法通过学习弱分类器,利用查询特征预测给定训练查询的检索性能增益。具有高性能增益的查询被选择用于下面的转换过程,以创建用于学习排名算法的伪标签。在标准LETOR数据集上的实验结果表明,我们提出的方法优于强基线。
Semi-supervised learning is a machine learning paradigm that can be applied to create pseudo labels from unlabeled data for learning a ranking model, when there is only limited or no training examples available. However, the effectiveness of semi-supervised learning in information retrieval (IR) can be hindered by the low quality pseudo labels, hence the need for the training query filtering that removes the low quality queries. In this paper, we assume two application scenarios with respect to the availability of human labels. First, for applications without any labeled data available, a clustering-based approach is proposed to select the high quality training queries. This approach selects the training queries following the empirical observation that the relevant documents of high quality training queries are highly coherent. Second, for applications with limited labeled data available, a classification-based approach is proposed. This approach learns a weak classifier to predict the retrieval performance gain of a given training query by making use of query features. The queries with high performance gains are selected for the following transduction process to create the pseudo labels for learning to rank algorithms. Experimental results on the standard LETOR dataset show that our proposed approaches outperform the strong baselines.
DOI: 10.1109/icdm.2006.22
发表时间: 2006-12
期刊: Sixth International Conference on Data Mining (ICDM'06)
影响因子: --
作者:
Xiangji Huang;Y. Huang;M. Wen;Aijun An;Y. Liu;Josiah Poon
通讯作者: Xiangji Huang;Y. Huang;M. Wen;Aijun An;Y. Liu;Josiah Poon
DOI: 10.1137/1.9781611972771.28
发表时间: 2007
期刊: --
影响因子: --
作者:
Hamed Valizadegan;P. Tan
通讯作者: Hamed Valizadegan;P. Tan
DOI: --
发表时间: 2007
影响因子: --
作者:
Tie-Yan Liu;Jun Xu;Tao Qin;Wen-Ying Xiong;Hang Li
通讯作者: Tie-Yan Liu;Jun Xu;Tao Qin;Wen-Ying Xiong;Hang Li
DOI: 10.1016/j.knosys.2013.01.032
发表时间: 2013-05-01
影响因子: 8.8
作者:
Leng, Yan;Xu, Xinyan;Qi, Guanghui
通讯作者: Qi, Guanghui
DOI: 10.1145/2396761.2398543
发表时间: 2012-10
期刊: Proceedings of the 21st ACM international conference on Information and knowledge management
影响因子: --
作者:
Xin Zhang;Ben He;Tiejian Luo;Baobin Li
通讯作者: Xin Zhang;Ben He;Tiejian Luo;Baobin Li