Learning to Automate Follow-up Question Generation using Process Knowledge for Depression Triage on Reddit Posts

Learning to Automate Follow-up Question Generation using Process Knowledge for Depression Triage on Reddit Posts
复制标题

DOI:
10.18653/v1/2022.clpsych-1.12
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Shrey Gupta;Anmol Agarwal;Manas Gaur;Kaushik Roy;Vignesh Narayanan;P. Kumaraguru;Amit P. Sheth
Shrey Gupta;Anmol Agarwal;Manas Gaur;Kaushik Roy;Vignesh Narayanan;P. Kumaraguru;Amit P. Sheth
中科院分区:
其他
文献类型:
--
作者:
Shrey Gupta;Anmol Agarwal;Manas Gaur;Kaushik Roy;Vignesh Narayanan;P. Kumaraguru;Amit P. Sheth

文献摘要

被引文献

相似文献

配备深度语言模型(DLM)的对话智能体(CA)在心理健康领域展现出巨大的潜力。显著的是,对话智能体已被用于为患者提供信息或治疗服务(例如认知行为疗法)。然而,在现有研究中,对话智能体在协助心理健康分诊方面的效用尚未得到探索,因为这需要有控制地生成后续问题(FQ),而在临床环境中,后续问题通常由心理健康专业人员(MHP)发起和引导。在“抑郁症”的背景下,我们的实验表明,与没有过程知识支持的深度语言模型相比,结合心理健康问卷中的过程知识的深度语言模型根据与PHQ - 9数据集中问题的相似度和最长公共子序列匹配,分别能生成更好的后续问题,其比例分别为12.54%和9.37%。尽管结合了过程知识,我们发现深度语言模型仍然容易产生幻觉,即生成冗余、不相关和不安全的后续问题。我们展示了使用现有数据集训练深度语言模型以生成符合临床过程知识的后续问题所面临的挑战。为了解决这一局限,我们与心理健康专业人员合作准备了一个基于扩展的PHQ - 9的数据集PRIMATE。PRIMATE包含关于PHQ - 9数据集中的特定问题是否已在用户对心理健康状况的初始描述中得到回答的注释。我们在有监督的环境下使用PRIMATE训练一个深度语言模型,以识别PHQ - 9中的哪些问题可以直接从用户的帖子中得到回答,哪些问题需要用户提供更多信息。通过基于马修斯相关系数(MCC)分数的性能分析,我们表明PRIMATE适合用于识别PHQ - 9中的问题,这些问题可以引导生成式深度语言模型朝着适合协助分诊的有控制的后续问题生成(且幻觉最小)。作为本研究一部分创建的数据集可从https://github.com/primate - mh/Primate2022获取。
Conversational Agents (CAs) powered with deep language models (DLMs) have shown tremendous promise in the domain of mental health. Prominently, the CAs have been used to provide informational or therapeutic services (e.g., cognitive behavioral therapy) to patients. However, the utility of CAs to assist in mental health triaging has not been explored in the existing work as it requires a controlled generation of follow-up questions (FQs), which are often initiated and guided by the mental health professionals (MHPs) in clinical settings. In the context of ‘depression’, our experiments show that DLMs coupled with process knowledge in a mental health questionnaire generate 12.54% and 9.37% better FQs based on similarity and longest common subsequence matches to questions in the PHQ-9 dataset respectively, when compared with DLMs without process knowledge support.Despite coupling with process knowledge, we find that DLMs are still prone to hallucination, i.e., generating redundant, irrelevant, and unsafe FQs. We demonstrate the challenge of using existing datasets to train a DLM for generating FQs that adhere to clinical process knowledge. To address this limitation, we prepared an extended PHQ-9 based dataset, PRIMATE, in collaboration with MHPs. PRIMATE contains annotations regarding whether a particular question in the PHQ-9 dataset has already been answered in the user’s initial description of the mental health condition. We used PRIMATE to train a DLM in a supervised setting to identify which of the PHQ-9 questions can be answered directly from the user’s post and which ones would require more information from the user. Using performance analysis based on MCC scores, we show that PRIMATE is appropriate for identifying questions in PHQ-9 that could guide generative DLMs towards controlled FQ generation (with minimal hallucination) suitable for aiding triaging. The dataset created as a part of this research can be obtained from https://github.com/primate-mh/Primate2022