Sequential Recommendation via Stochastic Self-Attention

Sequential Recommendation via Stochastic Self-Attention
复制标题

DOI:
10.1145/3485447.3512077
复制
发表时间:
2022-01
期刊:
Proceedings of the ACM Web Conference 2022
影响因子:
--
通讯作者:
Ziwei Fan;Zhiwei Liu;Yu Wang;Alice Wang;Zahra Nazari;Lei Zheng;Hao Peng;Philip S. Yu
Ziwei Fan;Zhiwei Liu;Yu Wang;Alice Wang;Zahra Nazari;Lei Zheng;Hao Peng;Philip S. Yu
中科院分区:
其他
文献类型:
--
作者:
Ziwei Fan;Zhiwei Liu;Yu Wang;Alice Wang;Zahra Nazari;Lei Zheng;Hao Peng;Philip S. Yu

文献摘要

被引文献

相似文献

序贯推荐是通过对用户先前行为的动态建模来预测下一个项目的一种方法,受到了广泛的关注。基于转换器的方法,它嵌入项目作为向量,并使用点积自我注意来衡量项目之间的关系,表现出上级的能力,现有的顺序方法。然而,用户的真实世界的顺序行为是不确定的,而不是确定的,提出了一个重大的挑战,目前的技术。我们进一步建议,基于点积的方法不能完全捕获协作传递性,这可以在序列内部的项目-项目转换中导出,并且对冷启动项目是有益的。我们进一步指出,BPR损失没有约束的积极和抽样的负面项目,这误导了优化。我们提出了一种新的随机自我注意(STOSA),以克服这些问题。特别是,STOSA将每个项目嵌入为随机高斯分布,其协方差编码不确定性。我们设计了一个新的Wasserstein自我注意模块来描述序列中的项目-项目位置关系,它有效地将不确定性纳入模型训练。沃瑟斯坦的关注也启发了协作传递性学习,因为它满足三角不等式。此外,我们引入了一个新的正则化项的排名损失,这保证了积极和消极的项目之间的相异性。在五个真实世界的基准数据集上进行的大量实验表明,该模型优于最先进的基线,特别是在冷启动项目上。该代码可在https://github.com/zfan20/STOSA上找到。
Sequential recommendation models the dynamics of a user’s previous behaviors in order to forecast the next item, and has drawn a lot of attention. Transformer-based approaches, which embed items as vectors and use dot-product self-attention to measure the relationship between items, demonstrate superior capabilities among existing sequential methods. However, users’ real-world sequential behaviors are uncertain rather than deterministic, posing a significant challenge to present techniques. We further suggest that dot-product-based approaches cannot fully capture collaborative transitivity, which can be derived in item-item transitions inside sequences and is beneficial for cold start items. We further argue that BPR loss has no constraint on positive and sampled negative items, which misleads the optimization. We propose a novel STOchastic Self-Attention (STOSA) to overcome these issues. STOSA, in particular, embeds each item as a stochastic Gaussian distribution, the covariance of which encodes the uncertainty. We devise a novel Wasserstein Self-Attention module to characterize item-item position-wise relationships in sequences, which effectively incorporates uncertainty into model training. Wasserstein attentions also enlighten the collaborative transitivity learning as it satisfies triangle inequality. Moreover, we introduce a novel regularization term to the ranking loss, which assures the dissimilarity between positive and the negative items. Extensive experiments on five real-world benchmark datasets demonstrate the superiority of the proposed model over state-of-the-art baselines, especially on cold start items. The code is available in https://github.com/zfan20/STOSA.