Confli-T5: An AutoPrompt Pipeline for Conflict Related Text Augmentation

Confli-T5: An AutoPrompt Pipeline for Conflict Related Text Augmentation
复制标题

DOI:
10.1109/bigdata55660.2022.10020509
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Eric Parolin;Yibo Hu;Latif Khan;Patrick T. Brandt;Javier Osorio;Vito D'Orazio
Eric Parolin;Yibo Hu;Latif Khan;Patrick T. Brandt;Javier Osorio;Vito D'Orazio
中科院分区:
其他
文献类型:
--
作者:
Eric Parolin;Yibo Hu;Latif Khan;Patrick T. Brandt;Javier Osorio;Vito D'Orazio

文献摘要

相似文献

自然语言处理(NLP)和大数据技术的最新进展对于科学家分析政治动荡和暴力、防止危害和促进全球冲突管理至关重要。政府机构和公安组织在基于深度学习的应用程序方面投入了大量资金,以研究全球冲突和政治暴力。然而,这种涉及文本分类、信息提取和其他与NLP相关的任务的应用程序需要大量的人工工作来注释/标记文本。虽然有限的标记数据可能会极大地损害模型的性能(过度拟合),但对注释任务的大量需求可能会使现实世界的应用变得不可行。为了解决这个问题,我们提出了Confli-T5,这是一种基于提示的方法,它利用现有政治学本体中的领域知识来生成冲突和调解领域中合成的但真实的标签文本样本。我们的模型允许从头开始生成文本数据,并使用我们新颖的双随机抽样机制来提高生成的样本的质量(一致性和一致性)。我们在六个与政治学研究相关的标准数据集上进行了实验,以显示Confli-T5的优越性。我们的代码是公开可用的1。
Recent advances in natural language processing (NLP) and Big Data technologies have been crucial for scientists to analyze political unrest and violence, prevent harm, and promote global conflict management. Government agencies and public security organizations have invested heavily in deep learning-based applications to study global conflicts and political violence. However, such applications involving text classification, information extraction, and other NLP-related tasks require extensive human efforts in annotating/labeling texts. While limited labeled data may drastically hurt the models’ performance (over-fitting), large demands on annotation tasks may turn real-world applications impracticable. To address this problem, we propose Confli-T5, a prompt-based method that leverages the domain knowledge from existing political science ontology to generate synthetic but realistic labeled text samples in the conflict and mediation domain. Our model allows generating textual data from the ground up and employs our novel Double Random Sampling mechanism to improve the quality (coherency and consistency) of the generated samples. We conduct experiments over six standard datasets relevant to political science studies to show the superiority of Confli-T5. Our codes are publicly available 1.