Discovering Communities with Self-Adaptive k Clustering in Microblog Data

Discovering Communities with Self-Adaptive k Clustering in Microblog Data
复制标题

DOI:
10.1109/cgc.2012.92
复制
发表时间:
2012-11
期刊:
2012 Second International Conference on Cloud and Green Computing
影响因子:
--
通讯作者:
Huang Ting;Dunlu Peng;Lidong Cao
Huang Ting;Dunlu Peng;Lidong Cao
中科院分区:
其他
文献类型:
--
作者:
Huang Ting;Dunlu Peng;Lidong Cao

文献摘要

被引文献

相似文献

如今,微博已经成为一种流行的社交网络服务,其人口在过去几年中令人难以置信地增加。许多商业公司将微博服务视为直接获取客户和潜在客户及时意见的不可或缺的媒介。社会网络中的社区是指一群具有相似兴趣或关注相同事物的人。微博社交网络服务中的用户社区识别对于识别热点话题或用户兴趣非常重要,有助于企业改进营销策略。然而,海量的非结构化tweet数据具有信息量大、涉及领域广、长度短、非结构化等特点,给有效挖掘其中隐藏的有价值社区带来了巨大的挑战。这使得tweets与传统的文本文档大不相同。为了更有效地分析数据,本文提出了一套技术来预处理推文,如词识别,类别匹配和数据标准化。提出了一种无监督学习的方法来自动将微博用户聚类到不同的社区。该方法根据微博数据的特点,开发了优化的CLARANS算法。在聚类过程中,还利用了推文之间的交互关系来提高聚类质量。此外,自适应k策略,使所提出的方法更适用。为了从多个方面考察该方法的性能,我们对新浪微博的微博数据进行了一系列实验。
Nowadays, microblogging has been a popular social network service whose population has incredibly increased in past few years. Many business companies regard microblogging service as an indispensable medium to directly obtain timely opinions from customers and potential customers. A community in social network refers to a crowd of people having similar interests or paying their attention on same things. User community recognition in microblogging social network service is very important for identifying hot topics or users' interests which are very helpful for companies to improve their marketing strategies. However, the massive non-structural tweet data brings tremendous challenge for efficiently mining the valuable communities hidden in it. Tweet data is characterized as containing massive information, being involved in large fields, short-length and non-structure. This makes tweets quite different from the conventional text documents. In order to analyze the data more effectively, in this paper, we propose a set of techniques to preprocess tweets, such as word identification, categories matching and data standardization. An unsupervised learning method has been presented to automatically cluster microblog users into different communities. In the method, an optimized CLARANS algorithm has been developed according to the characteristics of microblog data. During the process of clustering, the interactive relationship between tweets is also exploited to improve the clustering quality. In addition, a self-adaptive k strategy is employed to make the proposed approach more applicable. In order to investigate the performance of our approach from different aspects, we conducted a series of experiments with the microblog data collected from SINA Weibo.