Text Generation to Aid Depression Detection: A Comparative Study of Conditional Sequence Generative Adversarial Networks

Text Generation to Aid Depression Detection: A Comparative Study of Conditional Sequence Generative Adversarial Networks
复制标题

DOI:
10.1109/bigdata55660.2022.10020224
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
M. L. Tlachac;Walter Gerych;Kratika Agrawal;Benjamin R. Litterer;Nicholas Jurovich;Saitheeraj Thatigotla-Saitheeraj-That
M. L. Tlachac;Walter Gerych;Kratika Agrawal;Benjamin R. Litterer;Nicholas Jurovich;Saitheeraj Thatigotla-Saitheeraj-That
中科院分区:
其他
文献类型:
--
作者:
M. L. Tlachac;Walter Gerych;Kratika Agrawal;Benjamin R. Litterer;Nicholas Jurovich;Saitheeraj Thatigotla-Saitheeraj-That

文献摘要

相似文献

非结构化文本数据的语料库,如个人之间的短信,通常可以预测抑郁症等医疗问题。医疗保健应用程序中通常使用的文本数据具有很高的价值和多样性,但通常体积很小。生成标记的非结构化文本数据对于通过增加这些小数据集来改进模型以及促进匿名化非常重要。虽然存在标记数据生成方法,但并非所有方法都能很好地泛化到小数据集。因此,在这项工作中,我们对条件文本生成模型进行了非常必要的系统比较,这些模型由于其统一的架构而对小数据集很有希望。我们识别并实现了一组9个条件序列生成对抗网络用于文本生成,我们将其统称为cSeqGAN模型。这些模型具有两个正交设计维度:加权策略和反馈机制。我们进行了一项比较研究,评估了九种cSeqGAN模型在三种不同的带有抑郁和情绪标签的文本数据集上的生成能力。为了评估生成文本的质量和真实感,我们使用标准的机器学习指标以及通过用户研究进行的人类评估。而非条件模型产生预测文本,cSeqGAN模型产生更现实的文本。我们的比较研究奠定了坚实的基础,并为进一步的文本生成研究提供了重要的见解,特别是对于医疗保健领域常见的小数据集。
Corpuses of unstructured textual data, such as text messages between individuals, are often predictive of medical issues such as depression. The text data usually used in healthcare applications has high value and great variety, but is typically small in volume. Generating labeled unstructured text data is important to improve models by augmenting these small datasets, as well as to facilitate anonymization. While methods for labeled data generation exist, not all of them generalize well to small datasets. In this work, we thus perform a much needed systematic comparison of conditional text generation models that are promising for small datasets due to their unified architectures. We identify and implement a family of nine conditional sequence generative adversarial networks for text generation, which we collectively refer to as cSeqGAN models. These models are characterized along two orthogonal design dimensions: weighting strategies and feedback mechanisms. We conduct a comparative study evaluating the generation ability of the nine cSeqGAN models on three diverse text datasets with depression and sentiment labels. To assess the quality and realism of the generated text, we use standard machine learning metrics as well as human assessment via a user study. While the unconditioned models produced predictive text, the cSeqGAN models produced more realistic text. Our comparative study lays a solid foundation and provides important insights for further text generation research, particularly for the small datasets common within the healthcare domain.