Automated Construction of Lexicons to Improve Depression Screening With Text Messages

Automated Construction of Lexicons to Improve Depression Screening With Text Messages
复制标题

DOI:
10.1109/jbhi.2022.3203345
复制
发表时间:
2022-08
影响因子:
7.7
通讯作者:
M. L. Tlachac;Avantika Shrestha;M. Shah;Benjamin R. Litterer;Elke A. Rundensteiner
M. L. Tlachac;Avantika Shrestha;M. Shah;Benjamin R. Litterer;Elke A. Rundensteiner
中科院分区:
工程技术1区
文献类型:
--
作者:
M. L. Tlachac;Avantika Shrestha;M. Shah;Benjamin R. Litterer;Elke A. Rundensteiner

文献摘要

相似文献

鉴于抑郁症是最普遍的精神疾病之一,开发有效和不引人注目的诊断工具非常重要。最近的工作,屏幕抑郁症的短信利用模型依赖于词汇类别功能。鉴于文本消息的口语性质,这些模型的性能可能会受到正式词汇的限制。因此,我们提出了一个策略,自动构建替代词典,包含更多的相关和口语化的条款。具体来说,我们从小说,论坛和新闻语料库生成36个词汇。这些词汇,然后用于提取词汇类别功能的文本消息。我们利用机器学习模型来比较这些词汇类别特征的抑郁筛选能力。在我们的36个构建的词汇,14个取得了统计上显着较高的平均F1分数比预先存在的正式词汇和基本的词袋的方法。与现有的词典相比,我们的最佳表现词典将平均F1分数提高了10%。因此,我们证实了我们的假设,不太正式的词汇可以提高分类模型的性能,筛选抑郁症的短信。通过提供我们自动构建的词典,我们可以帮助未来的机器学习研究利用不太正式的文本。
Given that depression is one of the most prevalent mental illnesses, developing effective and unobtrusive diagnosis tools is of great importance. Recent work that screens for depression with text messages leverage models relying on lexical category features. Given the colloquial nature of text messages, the performance of these models may be limited by formal lexicons. We thus propose a strategy to automatically construct alternative lexicons that contain more relevant and colloquial terms. Specifically, we generate 36 lexicons from fiction, forum, and news corpuses. These lexicons are then used to extract lexical category features from the text messages. We utilize machine learning models to compare the depression screening capabilities of these lexical category features. Out of our 36 constructed lexicons, 14 achieved statistically significantly higher average F1 scores over the pre-existing formal lexicon and basic bag-of-words approach. In comparison to the pre-existing lexicon, our best performing lexicon increased the average F1 scores by 10%. We thus confirm our hypothesis that less formal lexicons can improve the performance of classification models that screen for depression with text messages. By providing our automatically constructed lexicons, we aid future machine learning research that leverages less formal text.