Understanding Depressive Symptoms and Psychosocial Stressors on Twitter: A Corpus-Based Study.

Understanding Depressive Symptoms and Psychosocial Stressors on Twitter: A Corpus-Based Study.
复制标题

DOI:
10.2196/jmir.6895
复制
发表时间:
2017-02-28
影响因子:
7.4
通讯作者:
Conway M
Conway M
中科院分区:
医学2区
文献类型:
--
作者:
Mowery D;Smith H;Cheney T;Stoddard G;Coppersmith G;Bryan C;Conway M

文献摘要

被引文献

相似文献

终身患病率为16.2%,是美国疾病负担的第五大贡献者。本研究的目的是建立在以前定性分析抑郁相关Twitter数据的基础上,描述一个全面的注释方案的发展(即编码方案),用于用《精神疾病诊断与统计手册》第5版(DSM 5)重度抑郁症状手动注释Twitter数据(例如,抑郁情绪、体重变化、精神运动性激动或发育迟缓)和《精神疾病诊断和统计手册》第四版(DSM-IV)心理社会压力源(例如,教育问题,与主要支持小组的问题,住房问题)。使用这种注释方案,我们开发了一个注释语料库,抑郁症状和心理社会压力获得性抑郁症,SAD语料库,包括9300推文随机抽样从Twitter应用程序编程接口(API)使用抑郁症相关的关键字(例如,抑郁,沮丧,悲伤)。我们的注释语料库的分析产生了几个关键的结果。首先,72.09%(6829/9473)的tweet包含相关的关键词,但并不表明有抑郁症状(例如,“我们正在经历一场新的经济萧条”)。其次,我们数据集中最普遍的症状是情绪低落和疲劳或能量损失。第三,不到2%的推文包含一个以上的抑郁相关类别(例如,思考或集中注意力的能力下降,抑郁情绪)。最后,我们发现在我们的注释数据集中,一些抑郁相关症状之间存在非常高的正相关性(例如,疲劳或能量损失与教育问题;教育问题与思考能力下降)。我们成功地开发了一个注释方案和注释语料库,SAD语料库,包括9300推随机选择的Twitter应用程序编程接口使用抑郁症相关的关键字。我们的分析表明,单独的关键字查询可能不适合公共卫生监测,因为上下文可以改变声明中关键字的含义。然而,后处理方法可能有助于减少噪音并改善使用社交媒体检测抑郁症状所需的信号。
With a lifetime prevalence of 16.2%, major depressive disorder is the fifth biggest contributor to the disease burden in the United States. The aim of this study, building on previous work qualitatively analyzing depression-related Twitter data, was to describe the development of a comprehensive annotation scheme (ie, coding scheme) for manually annotating Twitter data with Diagnostic and Statistical Manual of Mental Disorders, Edition 5 (DSM 5) major depressive symptoms (eg, depressed mood, weight change, psychomotor agitation, or retardation) and Diagnostic and Statistical Manual of Mental Disorders, Edition IV (DSM-IV) psychosocial stressors (eg, educational problems, problems with primary support group, housing problems). Using this annotation scheme, we developed an annotated corpus, Depressive Symptom and Psychosocial Stressors Acquired Depression, the SAD corpus, consisting of 9300 tweets randomly sampled from the Twitter application programming interface (API) using depression-related keywords (eg, depressed, gloomy, grief). An analysis of our annotated corpus yielded several key results. First, 72.09% (6829/9473) of tweets containing relevant keywords were nonindicative of depressive symptoms (eg, “we’re in for a new economic depression”). Second, the most prevalent symptoms in our dataset were depressed mood and fatigue or loss of energy. Third, less than 2% of tweets contained more than one depression related category (eg, diminished ability to think or concentrate, depressed mood). Finally, we found very high positive correlations between some depression-related symptoms in our annotated dataset (eg, fatigue or loss of energy and educational problems; educational problems and diminished ability to think). We successfully developed an annotation scheme and an annotated corpus, the SAD corpus, consisting of 9300 tweets randomly-selected from the Twitter application programming interface using depression-related keywords. Our analyses suggest that keyword queries alone might not be suitable for public health monitoring because context can change the meaning of keyword in a statement. However, postprocessing approaches could be useful for reducing the noise and improving the signal needed to detect depression symptoms using social media.