Extracting hierarchical structure of content groups from different social media platforms using multiple social metadata

Extracting hierarchical structure of content groups from different social media platforms using multiple social metadata
复制标题

DOI:
10.1007/s11042-017-4717-7
复制
发表时间:
2017-05
影响因子:
3.6
通讯作者:
Daichi Takehara;Ryosuke Harakawa;Takahiro Ogawa;M. Haseyama
Daichi Takehara;Ryosuke Harakawa;Takahiro Ogawa;M. Haseyama
中科院分区:
计算机科学4区
文献类型:
--
作者:
Daichi Takehara;Ryosuke Harakawa;Takahiro Ogawa;M. Haseyama

文献摘要

相似文献

一种用于检索用户所需内容的新方案,即,本文介绍了来自多个社交媒体平台的用户感兴趣的主题内容。在现有的检索方案中,用户首先选择一个特定的平台,然后输入查询到搜索引擎。如果用户不为他们的信息需求指定合适的平台,并且不输入与期望的内容相对应的合适的查询,则用户检索期望的内容变得困难。所提出的方案提取的层次结构的内容组(具有类似的主题的内容集)从不同的社交媒体平台,因此,它成为可行的检索所需的内容,即使用户不指定合适的平台,不输入合适的查询。本文的主要贡献有两个方面:(1)提出了一种新的特征提取方法--多社会元数据局部保持典型相关分析(LPCCA-MSM)方法,该方法能够检测出不受不同社会媒体平台限制的内容组。LPCCA-MSM使用多个社交元数据作为辅助信息,这与仅使用诸如文本或视觉特征的基于内容的信息的传统方法不同。(2)所提出的新检索方案可以实现来自不同社交媒体平台的分层内容结构化。提取的层次结构显示了内容组的各种抽象级别及其层次关系,这可以帮助用户选择与输入查询相关的主题。据我们所知,这种应用还没有进行深入的研究,因此,本文具有很强的新奇。为了验证上述贡献的有效性,对包含YouTube视频和维基百科文章的真实世界数据集进行了广泛的实验。
A novel scheme for retrieving users’ desired contents, i.e., contents with topics in which users are interested, from multiple social media platforms is presented in this paper. In existing retrieval schemes, users first select a particular platform and then input a query into the search engine. If users do not specify suitable platforms for their information needs and do not input suitable queries corresponding to the desired contents, it becomes difficult for users to retrieve the desired contents. The proposed scheme extracts the hierarchical structure of content groups (sets of contents with similar topics) from different social media platforms, and it thus becomes feasible to retrieve desired contents even if users do not specify suitable platforms and do not input suitable queries. This paper has two contributions: (1) A new feature extraction method, Locality Preserving Canonical Correlation Analysis with multiple social metadata (LPCCA-MSM) that can detect content groups without the boundaries of different social media platforms is presented in this paper. LPCCA-MSM uses multiple social metadata as auxiliary information unlike conventional methods that only use content-based information such as textual or visual features. (2) The proposed novel retrieval scheme can realize hierarchical content structuralization from different social media platforms. The extracted hierarchical structure shows various abstraction levels of content groups and their hierarchical relationships, which can help users select topics related to the input query. To the best of our knowledge, an intensive study on such an application has not been conducted; therefore, this paper has strong novelty. To verify the effectiveness of the above contributions, extensive experiments for real-world datasets containing YouTube videos and Wikipedia articles were conducted.