Sentiment Analysis on Social Media Content

Sentiment Analysis on Social Media Content
复制标题

社交媒体内容的情感分析

DOI:
--
复制
发表时间:
2020
期刊:
arXiv.org
影响因子:
--
通讯作者:
John Mcgonical
John Mcgonical
中科院分区:
--
文献类型:
--
作者:
Antony Samuels;John Mcgonical

文献摘要

被引文献

相似文献

如今,来自世界各地的人们使用社交媒体网站分享信息。例如,Twitter是一个平台,用户可以在其中发送、阅读被称为tweets的帖子,并与不同的社区进行互动。用户分享他们的日常生活,发布他们对品牌和地点等一切的看法。公司可以通过收集与他们的意见相关的数据,从这个庞大的平台中受益。本文的目的是提出一个模型,可以执行情感分析的真实的数据从Twitter上收集。Twitter中的数据是高度非结构化的,这使得它很难分析。然而,我们提出的模型与该领域的先前工作不同,因为它结合了监督和无监督机器学习算法的使用。执行情感分析的过程如下:直接从Twitter API中提取推文,然后执行数据的清理和发现。之后,数据被输入到几个模型中进行训练。每一条推文都根据其情绪进行分类,无论是积极的,消极的还是中立的。数据收集了两个主题麦当劳和肯德基,以显示哪个餐厅更受欢迎。使用了不同的机器学习算法。使用交叉验证和f-score等各种测试指标对这些模型的结果进行了测试。此外,我们的模型在直接从Twitter中提取的挖掘文本上表现出很强的性能。
Nowadays, people from all around the world use social media sites to share information. Twitter for example is a platform in which users send, read posts known as tweets and interact with different communities. Users share their daily lives, post their opinions on everything such as brands and places. Companies can benefit from this massive platform by collecting data related to opinions on them. The aim of this paper is to present a model that can perform sentiment analysis of real data collected from Twitter. Data in Twitter is highly unstructured which makes it difficult to analyze. However, our proposed model is different from prior work in this field because it combined the use of supervised and unsupervised machine learning algorithms. The process of performing sentiment analysis as follows: Tweet extracted directly from Twitter API, then cleaning and discovery of data performed. After that the data were fed into several models for the purpose of training. Each tweet extracted classified based on its sentiment whether it is a positive, negative or neutral. Data were collected on two subjects McDonalds and KFC to show which restaurant has more popularity. Different machine learning algorithms were used. The result from these models were tested using various testing metrics like cross validation and f-score. Moreover, our model demonstrates strong performance on mining texts extracted directly from Twitter.