Multi-kernel SVM based depression recognition using social media data

Multi-kernel SVM based depression recognition using social media data
复制标题

DOI:
10.1007/s13042-017-0697-1
复制
发表时间:
2019-01-01
影响因子:
5.6
通讯作者:
Dang, Jianwu
Dang, Jianwu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Peng, Zhichao;Hu, Qinghua;Dang, Jianwu

文献摘要

被引文献

相似文献

抑郁症已成为世界第四大疾病。然而,与高发病率相比,抑郁症的就医率却很低,这是由于精神问题的诊断困难。社交媒体打开了一个评估用户心理状态的窗口。随着互联网的迅速发展,人们习惯于通过社交媒体表达自己的想法和感受。因此,社交媒体提供了一种新的方式来发现潜在的抑郁症患者。在本文中,我们提出了一个多核支持向量机模型识别抑郁的人。从用户的社交媒体中提取三类特征,即用户微博文本、用户个人资料和用户行为来描述用户的情况。针对社交媒体语言的新特点,构建了由文本情感词典和表情符号词典组成的微博情感词典,提取微博文本特征进行词频统计。考虑到文本特征和其他两个特征之间的异质性,我们采用多核支持向量机的方法,自适应地选择最佳的内核为不同的功能,发现可能患有抑郁症的用户。与朴素贝叶斯、决策树、KNN、单核支持向量机和集成方法(libD 3C)相比,多核支持向量机方法识别抑郁症的错误率分别为38%、43%、22%、21%和11%,降低到16.54%。这表明多核SVM方法是基于社交媒体数据发现抑郁人群的最合适方法。
Depression has become the world's fourth major disease. Compared with the high incidence, however, the rate of depression medical treatment is very low because of the difficulty of diagnosis of mental problems. The social media opens one window to evaluate the users' mental status. With the rapid development of Internet, people are accustomed to express their thoughts and feelings through social media. Thus social media provides a new way to find out the potential depressed people. In this paper, we propose a multi-kernel SVM based model to recognize the depressed people. Three categories of features, user microblog text, user profile and user behaviors, are extracted from their social media to describe users' situations. According to the new characteristics of social media language, we build a special emotional dictionary consisted of text emotional dictionary and emoticon dictionary to extract microblog text features for word frequency statistics. Considering the heterogeneity between text feature and another two features, we employ multi-kernel SVM methods to adaptively select the optimal kernel for different features to find out users who may suffer from depression. Compared with Naive Bayes, Decision Trees, KNN, single-kernel SVM and ensemble method (libD3C), whose error reduction rates are 38, 43, 22, 21 and 11% respectively, the error rate of multi-kernel SVM method for identifying the depressed people is reduced to 16.54%. This indicates that the multi-kernel SVM method is the most appropriate way to find out depressed people based on social media data.