Semantic Weighted Multi-View Clustering for Web Content

Semantic Weighted Multi-View Clustering for Web Content
复制标题

Web 内容的语义加权多视图聚类

DOI:
10.1109/access.2019.2939334
复制
发表时间:
2019-09
期刊:
影响因子:
3.9
通讯作者:
Ma Zhiyi
Ma Zhiyi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Gong Xiaolong;Huang Linpeng;Luo Tiancheng;Ma Zhiyi

文献摘要

参考文献

相似文献

聚类是一个长期存在的重要研究问题。然而,当处理来自不同类型信息资源的大规模Web数据时,如用户个人资料,评论,用户偏好等,所有这些方面都可以被视为不同的视图,并且通常承认数据的相同底层聚类。本文提出了一种新的语义加权非负矩阵分解(<inline-formula><tex-math notation="LaTeX">$SWNMF$</tex-math></inline-formula>)多视图聚类框架,该框架提供了一种高效的加权矩阵分解框架,灵活地处理多视图Web内容,并易于探索数据语义空间中的稀疏性问题。具体而言,数据集的每个视图构成一个庞大的稀疏矩阵,导致矩阵分解过程中的非鲁棒性,进而影响聚类结果的准确性。为了解决上述问题,我们尝试使用用户给出的一些偏好信息(如评分值)作为潜在的语义信息来处理那些在每个数据点中未观察到的特征,从而解决所有视图矩阵中的稀疏性问题。为了在大型语料库中实现多视图的联合收割机组合,我们提出的<inline-formula><tex-math notation="LaTeX">$SWNMF$</tex-math></inline-formula>的总体目标是最小化<inline-formula><tex-math notation="LaTeX">l_{2,1}$</tex-math></inline-formula>-范数下的加权<italic>非负矩阵分解</italic>(NMF)的损失函数和<inline-formula><tex-math notation="LaTeX">$F$</tex-math></inline-formula>-范数下的共正则化约束。我们的大规模多视图Web数据集上的大量实验证明了我们的解决方案的竞争力的性能。
Clustering is a long-standing important research problem. However, it remains challenging when handling large-scale web data from different types of information resources such as user profile, comments, user preferences and so on. All these aspects can be seen as different views and often admit the same underlying clustering of the data. In this paper, we present a novel Semantic Weighted Non-negative Matrix Factorization (<inline-formula> <tex-math notation="LaTeX">$SWNMF$ </tex-math></inline-formula>) multi-view clustering framework, which can provide an efficient weighted matrix factorization framework, dexterously manipulate multi-view web content, and easily explore the sparseness problem in semantic space of data. Specifically, each view of dataset forming a huge sparse matrix, which results in the non-robust characteristic during the matrix decomposition process, and further influences the accuracy of clustering results. To address above problem, we attempt to use some preference information (e.g. rating values) given by the users as latent semantic information to handle those features that are unobserved in each data point so as to resolve the sparseness problem in all views matrices. To combine multiple views in our large corpus, the overall objective of our proposed <inline-formula> <tex-math notation="LaTeX">$SWNMF$ </tex-math></inline-formula> is to minimize the loss function of weighted <italic>non-negative matrix factorization</italic> (NMF) under the <inline-formula> <tex-math notation="LaTeX">$l_{2,1}$ </tex-math></inline-formula>-norm and the co-regularized constraint under the <inline-formula> <tex-math notation="LaTeX">$F$ </tex-math></inline-formula>-norm. Extensive experiments on our large-scale multi-view web datasets demonstrate the competitive performance of our solution.
DOI: --
发表时间: 2011-02
期刊: --
影响因子: --
作者:
Zeynep Akata;Christian Thurau;C. Bauckhage
通讯作者: Zeynep Akata;Christian Thurau;C. Bauckhage
DOI: 10.1137/1.9781611972795.5
发表时间: 2009-06
期刊: --
影响因子: --
作者:
Xinhai Liu;Shi Yu;Y. Moreau;B. Moor;W. Glänzel;Frizo A. L. Janssens
通讯作者: Xinhai Liu;Shi Yu;Y. Moreau;B. Moor;W. Glänzel;Frizo A. L. Janssens
DOI: 10.1145/2063576.2063621
发表时间: 2011-10
期刊: --
影响因子: --
作者:
Hua Wang;Heng Huang;C. Ding
通讯作者: Hua Wang;Heng Huang;C. Ding
DOI: --
发表时间: 2005-12
期刊: Lawrence Berkeley National Laboratory
影响因子: --
作者:
C. Ding;Xiaofeng He;H. Simon;Rong Jin
通讯作者: C. Ding;Xiaofeng He;H. Simon;Rong Jin
DOI: --
发表时间: --
期刊: --
影响因子: --
作者:
Chenping Hou;Changshui Zhang;Yi Wu;F. Nie
通讯作者: Chenping Hou;Changshui Zhang;Yi Wu;F. Nie