Complementary or Substitutive? A Novel Deep Learning Method to Leverage Text-image Interactions for Multimodal Review Helpfulness Prediction

Complementary or Substitutive? A Novel Deep Learning Method to Leverage Text-image Interactions for Multimodal Review Helpfulness Prediction
复制标题

DOI:
10.1016/j.eswa.2022.118138
复制
发表时间:
2022-07
期刊:
Expert Syst. Appl.
影响因子:
--
通讯作者:
S. Xiao;Gang Chen;Chenghong Zhang-;Xiangge Li
S. Xiao;Gang Chen;Chenghong Zhang-;Xiangge Li
中科院分区:
其他
文献类型:
--
作者:
S. Xiao;Gang Chen;Chenghong Zhang-;Xiangge Li

文献摘要

相似文献

随着移动互联网的蓬勃发展,多模态评论(即文本和图像评论)变得越来越普遍,并在客户决策中发挥着重要作用。然而,在进行多模态评论有用性预测(MRHP)时,由于文本和图像之间的信息交互而变得困难。评论文本(图像)中的信息可以补充或替代视觉(文本)评论信息。此外,在某些情况下,文本(图像)本身可能主要构成评论的诊断价值,而在其他情况下,它们可能被客户共同认为是有用的。在这项研究中,我们深入研究通过对文本-图像交互进行建模来进行 MRPH。我们提出了一种新颖的多模态深度学习方法,该方法利用文本和图像之间的互补和替代效应,并进一步协调它们以实现 MRHP。对大规模在线评论数据集的实证评估表明,我们提出的方法优于基准,表明其预测多模式评论有用性的强大能力。探索性分析为理解评论文本和图像之间的互补替代交互模式提供了见解。
With the flourishing of mobile Internet, the multimodal reviews (i.e., reviews with both texts and images) are becoming prevalent and playing an important role in customer decision makings. However, when making multimodal review helpfulness prediction (MRHP), it becomes difficult due to the information interaction between text and images. The information in review text (images) can be either complementary or substitutive to visual (textual) review information. Moreover, the text (images) itself may constitute the review’s diagnostic value predominantly in some cases, whereas they could be jointly perceived as useful by customers in others. In this study, we delve to conduct MRPH by modeling their text-image interactions. We proposed a novel multimodal deep learning method that exploits the complementation and substitution effects between text and images and further coordinates them for MRHP. Empirical evaluation on a large-scale online review dataset shows that our proposed method outperformed the benchmarks, indicating its powerful capability to predict the helpfulness of multimodal reviews. Exploratory analysis renders insights for understanding the complementary-substitutive interaction patterns between review text and images.