Scaling Archived Social Media Data Analysis Using a Hadoop Cloud

Scaling Archived Social Media Data Analysis Using a Hadoop Cloud
复制标题

DOI:
10.1109/cloud.2013.120
复制
发表时间:
2013-06
期刊:
2013 IEEE Sixth International Conference on Cloud Computing
影响因子:
--
通讯作者:
Javier Conejero;P. Burnap;O. Rana;Jeffrey Morgan
Javier Conejero;P. Burnap;O. Rana;Jeffrey Morgan
中科院分区:
其他
文献类型:
--
作者:
Javier Conejero;P. Burnap;O. Rana;Jeffrey Morgan

文献摘要

相似文献

近年来,人们对支持社交媒体分析进行营销、意见分析和理解社区凝聚力的兴趣日益浓厚。社交媒体数据符合“大数据”的许多分类——即数量、速度和多样性。一般来说,需要对大量数据进行有效和及时的分析。已经报道了各种计算基础设施来实现这一点。我们展示了支持Twitter数据情绪和张力分析的COSMOS平台,并演示了如何使用OpenNebula云环境和基于Hadoop的Map/ reduce分析来扩展这个平台。特别地,我们描述了从性能角度来看最有用的系统配置类型——例如,应该如何分布基础架构中的虚拟机以减少分析性能的可变性。我们使用由数百万条Twitter消息组成的数据集来演示该方法,并在两种类型的云基础设施上进行分析。
Over recent years, there has been an emerging interest in supporting social media analysis for marketing, opinion analysis and understanding community cohesion. Social media data conforms to many of the categorisations attributed to "big-data" -- i.e. volume, velocity and variety. Generally analysis needs to be undertaken over large volumes of data in an efficient and timely manner. A variety of computational infrastructures have been reported to achieve this. We present the COSMOS platform supporting sentiment and tension analysis on Twitter data, and demonstrate how this platform can be scaled using the OpenNebula Cloud environment with Map/Reduce-based analysis using Hadoop. In particular, we describe the types of system configurations that would be most useful from a performance perspective -- i.e. how virtual machines in the infrastructure should be distributed to reduce variability in the analysis performance. We demonstrate the approach using a data set consisting of several million Twitter messages, analysed over two types of Cloud infrastructure.