Monitoring and Recognizing Enterprise Public Opinion from High-Risk Users Based on User Portrait and Random Forest Algorithm

Monitoring and Recognizing Enterprise Public Opinion from High-Risk Users Based on User Portrait and Random Forest Algorithm
复制标题

DOI:
10.3390/axioms10020106
复制
发表时间:
2021-06-01
期刊:
影响因子:
2
通讯作者:
Cong, Guodong
Cong, Guodong
中科院分区:
数学3区
文献类型:
--
作者:
Chen, Tinggui;Yin, Xiaohua;Cong, Guodong

文献摘要

被引文献

相似文献

随着自媒体技术的快速发展,网民可以在网络平台上自由地表达对企业产品的看法。因此,网络上的企业舆情成为一个突出的问题。一些网民发表的负面评论可能会引发负面舆论,对企业的形象产生重大影响。本文从帮助企业应对负面舆论的角度出发,将用户画像技术与随机森林算法相结合,帮助企业识别那些发表过负面评论、可能引发负面舆论的高风险用户。通过这种方式,企业可以监控高风险用户的舆情,防止负面舆情事件的发生。首先,我们抓取参与产品体验讨论的用户信息,构建企业舆情用户画像。然后,将画像的特征量化为用户活跃度、用户影响力、用户情感倾向等指标,并对指标进行排序。根据指标排序,将用户分为高危、中危、低危三类。其次,基于随机森林算法,建立了该分类的监督高风险用户识别模型。然后,训练好的随机森林标识符可以用来预测新发布的舆情信息的作者是否是高风险用户。最后,采用反向传播神经网络算法对用户进行识别,并与模型识别结果进行比较。结果表明,反向传播神经网络的平均识别准确率仅为72.33%,而本文构建的模型的平均识别准确率高达98.49%,验证了所提出的随机森林识别方法的可行性和准确性。
With the rapid development of "We media" technology, netizens can freely express their opinions regarding enterprise products on a network platform. Consequently, online public opinion about enterprises has become a prominent issue. Negative comments posted by some netizens may trigger negative public opinion, which can have a significant impact on an enterprise's image. From the perspective of helping enterprises deal with negative public opinion, this paper combines user portrait technology and a random forest algorithm to help enterprises identify high-risk users who have posted negative comments and thus may trigger negative public opinion. In this way, enterprises can monitor the public opinion of high-risk users to prevent negative public opinion events. Firstly, we crawled the information of users participating in discussions of product experience, and we constructed a portrait of enterprise public opinion users. Then, the characteristics of the portraits were quantified into indicators such as the user's activity, the user's influence, and the user's emotional tendency, and the indicators were sorted. According to the order of the indicators, the users were divided into high-risk, moderate-risk, and low-risk categories. Next, a supervised high-risk user identification model for this classification was established, based on a random forest algorithm. In turn, the trained random forest identifier can be used to predict whether the authors of newly published public opinion information are high-risk users. Finally, a back propagation neural network algorithm was used to identify users and compared with the results of model recognition in this paper. The results showed that the average recognition accuracy of the back propagation neural network is only 72.33%, while the average recognition accuracy of the model constructed in this paper is as high as 98.49%, which verifies the feasibility and accuracy of the proposed random forest recognition method.