What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation

What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation
复制标题

DOI:
10.1145/3383313.3412249
复制
发表时间:
2020-07
期刊:
Proceedings of the 14th ACM Conference on Recommender Systems
影响因子:
--
通讯作者:
Gustavo Penha;C. Hauff
Gustavo Penha;C. Hauff
中科院分区:
其他
文献类型:
--
作者:
Gustavo Penha;C. Hauff

文献摘要

被引文献

相似文献

最近,BERT 等经过大量预训练的 Transformer 模型在语言建模方面表现出非常强大的能力,在众多下游任务中取得了令人印象深刻的结果。研究还表明,它们在预训练后将事实知识隐式存储在参数中。了解 LM 的预训练过程实际学到的内容是使用和改进对话推荐系统 (CRS) 的关键一步。我们首先研究现成的预训练 BERT 对书籍、电影和音乐等推荐项目“了解”多少。为了分析 BERT 参数中存储的知识,我们使用不同的探针(即检查有关某些属性的训练模型的任务),这些探针需要不同类型的知识来解决,即基于内容的知识和基于协作的知识。基于内容的知识是要求模型将项目标题与其内容信息(例如文本描述和流派)相匹配的知识。相比之下,基于协作的知识要求模型根据评级等社区互动来匹配具有相似项目的项目。我们借助 BERT 的掩码语言建模 (MLM) 头来探究其关于项目类型的知识,并带有完形填空风格的提示。此外,我们使用 BERT 的下一句预测 (NSP) 头和表示相似度 (SIM) 来比较相关和不相关的搜索和推荐查询文档输入,以探索 BERT 是否可以在不进行任何微调的情况下将相关项排名第一。最后,我们研究了 BERT 在会话推荐下游任务中的表现。为此,我们对 BERT 进行了微调,使其充当基于检索的 CRS。总的来说,我们的实验表明:(i)BERT 在其参数中存储了有关书籍、电影和音乐内容的知识; (ii) 其基于内容的知识多于基于协作的知识; (iii) 面对对抗性数据时,对话式推荐失败。
Heavily pre-trained transformer models such as BERT have recently shown to be remarkably powerful at language modelling, achieving impressive results on numerous downstream tasks. It has also been shown that they implicitly store factual knowledge in their parameters after pre-training. Understanding what the pre-training procedure of LMs actually learns is a crucial step for using and improving them for Conversational Recommender Systems (CRS). We first study how much off-the-shelf pre-trained BERT “knows” about recommendation items such as books, movies and music. In order to analyze the knowledge stored in BERT’s parameters, we use different probes (i.e., tasks to examine a trained model regarding certain properties) that require different types of knowledge to solve, namely content-based and collaborative-based. Content-based knowledge is knowledge that requires the model to match the titles of items with their content information, such as textual descriptions and genres. In contrast, collaborative-based knowledge requires the model to match items with similar ones, according to community interactions such as ratings. We resort to BERT’s Masked Language Modelling (MLM) head to probe its knowledge about the genre of items, with cloze style prompts. In addition, we employ BERT’s Next Sentence Prediction (NSP) head and representations’ similarity (SIM) to compare relevant and non-relevant search and recommendation query-document inputs to explore whether BERT can, without any fine-tuning, rank relevant items first. Finally, we study how BERT performs in a conversational recommendation downstream task. To this end, we fine-tune BERT to act as a retrieval-based CRS. Overall, our experiments show that: (i) BERT has knowledge stored in its parameters about the content of books, movies and music; (ii) it has more content-based knowledge than collaborative-based knowledge; and (iii) fails on conversational recommendation when faced with adversarial data.