Monitoring Depression Trend on Twitter during the COVID-19 Pandemic

Monitoring Depression Trend on Twitter during the COVID-19 Pandemic
复制标题

DOI:
10.2196/preprints.26769
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Yipeng Zhang;Hanjia Lyu;Yubao Liu;Xiyang Zhang;Yu Wang;Jiebo Luo
Yipeng Zhang;Hanjia Lyu;Yubao Liu;Xiyang Zhang;Yu Wang;Jiebo Luo
中科院分区:
其他
文献类型:
--
作者:
Yipeng Zhang;Hanjia Lyu;Yubao Liu;Xiyang Zhang;Yu Wang;Jiebo Luo

文献摘要

被引文献

相似文献

背景COVID-19疫情严重影响了人们的日常生活,并在全球范围内造成了巨大的经济损失。传闻证据表明,这一流行病增加了人口中的抑郁程度。然而,抑郁症的检测和监测过程中的抑郁症缺乏系统的研究。目的本研究旨在(1)开发一种方法,通过分析他们的推文来准确识别抑郁症患者;(2)监测Twitter上的人群抑郁水平。为了研究这个问题,我们设计了一个有效的正则表达式为基础的搜索方法,并创建了迄今为止最大的英文Twitter抑郁症数据集包含2,575个不同的识别抑郁症用户(N=2,575)与他们过去的推文。为了研究抑郁症对人们Twitter语言的影响,我们在数据集上训练了三个基于transformer的抑郁症分类模型,通过逐渐增加的训练规模来评估它们的性能,并比较模型的“tweet chunk”级别和用户级别的性能。此外,受心理学研究的启发,我们创建了一个融合分类器,将深度学习模型分数与心理文本特征和用户的人口统计信息相结合,并研究这些特征与抑郁信号的关系。最后,我们展示了我们的模型在COVID-19大流行期间的两个应用,以证明我们的模型能够监测群体水平和人群水平的抑郁趋势。结果我们的融合模型在包含446人(N=446)的测试集上显示出78.9%的准确率,其中一半被识别为患有抑郁症。尽责性,神经质,第一人称代词的出现,谈论生物过程,如吃饭和睡觉,谈论权力,表现出悲伤被证明是抑郁症分类的重要特征。此外,当用于监测抑郁症趋势时,我们的模型显示,抑郁症用户通常比对照组更晚根据他们的推文对大流行做出反应。它还表明,美国的三个州-纽约(NY),加州(CA)和佛罗里达(FL)-共享一个类似的抑郁症的趋势,作为整个美国人口。与纽约州和加利福尼亚州相比,佛罗里达州的人们表现出明显较低的抑郁水平。结论本研究提出了一种有效的方法,可用于分析不同人群的抑郁水平的Twitter。我们希望这项研究能够提高研究人员和公众对COVID-19对人们心理健康影响的认识。非侵入式监测系统还可以快速适应COVID-19以外的其他重大事件,并可能在未来爆发期间有用。
BACKGROUND The COVID-19 pandemic has severely affected people’s daily lives and caused tremendous economic loss worldwide. Anecdotal evidence suggests that the pandemic has increased the depression level among the population. However, systematic studies of depression detection and monitoring during the depression are lacking. OBJECTIVE This study aims (1) to develop a method to accurately identify people with depression by analyzing their tweets and (2) to monitor the population-wise depression level on Twitter. METHODS To study this subject, we design an effective regular expression-based search method and create by far the largest English Twitter depression dataset containing 2,575 distinct identified depression users (N=2,575) with their past tweets. To examine the effect of depression on people’s Twitter language, we train three transformer-based depression classification models on the dataset, evaluate their performance with progressively increased training sizes, and compare the model’s “tweet chunk”-level and user-level performances. Furthermore, inspired by psychological studies, we create a fusion classifier that combines deep learning model scores with psychological text features and users’ demographic information and investigate these features’ relations to depression signals. Finally, we demonstrate our model’s capability of monitoring both group-level and population-level depression trends by presenting two of its applications during the COVID-19 pandemic. RESULTS Our fusion model demonstrates an accuracy of 78.9% on a test set containing 446 people (N=446), half of which are identified as suffering from depression. Conscientiousness, neuroticism, appearance of first-person pronouns, talking about biological processes such as eat and sleep, talking about power, and exhibiting sadness are shown to be important features in depression classification. Further, when used for monitoring the depression trend, our model shows that depressive users, in general, respond to the pandemic later than the control group based on their tweets. It is also shown that three states of the United States - New York (NY), California (CA), and Florida (FL) - share a similar depression trend as the whole US population. When compared to NY and CA, people in FL demonstrate a significantly lower level of depression. CONCLUSIONS This study proposes an efficient method that can be used to analyze the depression level of different groups of people on Twitter. We hope this study can raise awareness among researchers and the general public of COVID-19’s impact on people’s mental health. The non-invasive monitoring system can also be rapidly adapted to other big events besides COVID-19 and might be useful during future outbreaks.