Balanced Knowledge Distillation with Contrastive Learning for Document Re-ranking

Balanced Knowledge Distillation with Contrastive Learning for Document Re-ranking
复制标题

DOI:
10.1145/3578337.3605120
复制
发表时间:
2023-08
期刊:
Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval
影响因子:
--
通讯作者:
Yingrui Yang;Shanxiu He;Yifan Qiao;Wentai Xie;Tao Yang
Yingrui Yang;Shanxiu He;Yifan Qiao;Wentai Xie;Tao Yang
中科院分区:
其他
文献类型:
--
作者:
Yingrui Yang;Shanxiu He;Yifan Qiao;Wentai Xie;Tao Yang

文献摘要

相似文献

知识蒸馏通常用于训练神经文档排序模型,通过雇用教师来指导模型细化。由于教师不可能在所有情况下都是正确的,学生和教师模型之间的过度校准可能会使培训效果降低。本文重点研究了KL发散损失用于文档重排序中的知识提取,并重新审视了知识提取与显式对比学习的平衡。所提出的损失函数在模仿教师的行为时采取保守的方法,并且允许学生有时通过训练偏离教师的模型。本文提出了分析结果与MS MARCO通道的评估,以验证的有用性,提出的损失为基础的变压器ColBERT重新排名。
Knowledge distillation is commonly used in training a neural document ranking model by employing a teacher to guide model refinement. As a teacher may not be correct in all cases, over-calibration between the student and teacher models can make training less effective. This paper focuses on the KL divergence loss used for knowledge distillation in document re-ranking, and re-visits balancing of knowledge distillation with explicit contrastive learning. The proposed loss function takes a conservative approach in imitating teacher's behavior, and allows student to deviate from a teacher's model sometimes through training. This paper presents analytic results with an evaluation on MS MARCO passages to validate the usefulness of the proposed loss for the transformer-based ColBERT re-ranking.