Balanced Knowledge Distillation with Contrastive Learning for Document Re-ranking
Balanced Knowledge Distillation with Contrastive Learning for Document Re-ranking
复制标题
DOI:
10.1145/3578337.3605120
复制
发表时间:
2023-08
期刊:
影响因子:
--
通讯作者:
Yingrui Yang;Shanxiu He;Yifan Qiao;Wentai Xie;Tao Yang
中科院分区:
文献类型:
--
作者:
Yingrui Yang;Shanxiu He;Yifan Qiao;Wentai Xie;Tao Yang
Knowledge distillation is commonly used in training a neural document ranking model by employing a teacher to guide model refinement. As a teacher may not be correct in all cases, over-calibration between the student and teacher models can make training less effective. This paper focuses on the KL divergence loss used for knowledge distillation in document re-ranking, and re-visits balancing of knowledge distillation with explicit contrastive learning. The proposed loss function takes a conservative approach in imitating teacher's behavior, and allows student to deviate from a teacher's model sometimes through training. This paper presents analytic results with an evaluation on MS MARCO passages to validate the usefulness of the proposed loss for the transformer-based ColBERT re-ranking.