A probabilistic graphical model for topic and preference discovery on social media

A probabilistic graphical model for topic and preference discovery on social media
复制标题

DOI:
10.1016/j.neucom.2011.05.039
复制
发表时间:
2012-10
期刊:
影响因子:
6
通讯作者:
Lu Liu;Feida Zhu;Lei Zhang;Shiqiang Yang
Lu Liu;Feida Zhu;Lei Zhang;Shiqiang Yang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Lu Liu;Feida Zhu;Lei Zhang;Shiqiang Yang

文献摘要

被引文献

相似文献

如今,许多 Web 应用程序的蓬勃发展都依赖于为大规模多媒体数据提供服务,例如用于照片的 Flickr 和用于视频的 YouTube。然而,这些数据虽然内容丰富,但文本描述信息通常很少。例如,视频剪辑通常仅与几个标签相关联。此外,文字描述往往过于具体于视频内容。这些特征使得在此类数据上以令人满意的粒度发现主题非常具有挑战性。在本文中,我们提出了一种名为偏好主题模型(PTM)的生成概率模型,引入用户偏好的维度来增强文本信息的不足。 PTM 是一个将用户偏好发现和文档主题挖掘任务结合在一起的统一框架。通过对用户-文档交互进行建模,PTM 不仅可以同时发现主题和偏好,还可以使它们在统一的框架中相互告知并受益。因此,PTM 可以从稀疏数据中提取更好的主题和偏好。在现实生活视频应用数据上的实验结果表明,在基于聚类的评估方面,PTM 在发现信息主题和偏好方面优于 LDA。此外,DBLP 数据上的实验结果表明,PTM 是一个通用模型,可以应用于其他类型的用户-文档交互。
Many web applications today thrive on offering services for large-scale multimedia data, e.g., Flickr for photos and YouTube for videos. However, these data, while rich in content, are usually sparse in textual descriptive information. For example, a video clip is often associated with only a few tags. Moreover, the textual descriptions are often overly specific to the video content. Such characteristics make it very challenging to discover topics at a satisfactory granularity on this kind of data. In this paper, we propose a generative probabilistic model named Preference-Topic Model (PTM) to introduce the dimension of user preferences to enhance the insufficient textual information. PTM is a unified framework to combine the tasks of user preference discovery and document topic mining together. Through modeling user-document interactions, PTM cannot only discover topics and preferences simultaneously, but also enable them to inform and benefit each other in a unified framework. As a result, PTM can extract better topics and preferences from sparse data. The experimental results on real-life video application data show that PTM is superior to LDA in discovering informative topics and preferences in terms of clustering-based evaluations. Furthermore, the experimental results on DBLP data demonstrate that PTM is a general model which can be applied to other kinds of user–document interactions.