Multi-CoPED: A Multilingual Multi-Task Approach for Coding Political Event Data on Conflict and Mediation Domain

Multi-CoPED: A Multilingual Multi-Task Approach for Coding Political Event Data on Conflict and Mediation Domain
复制标题

DOI:
10.1145/3514094.3534178
复制
发表时间:
2022-07
期刊:
Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society
影响因子:
--
通讯作者:
Erick Skorupa Parolin;Mohammad Javad Hosseini;Yibo Hu;Latif Khan;Patrick T. Brandt;Javier Osorio;Vito D'Orazio
Erick Skorupa Parolin;Mohammad Javad Hosseini;Yibo Hu;Latif Khan;Patrick T. Brandt;Javier Osorio;Vito D'Orazio
中科院分区:
其他
文献类型:
--
作者:
Erick Skorupa Parolin;Mohammad Javad Hosseini;Yibo Hu;Latif Khan;Patrick T. Brandt;Javier Osorio;Vito D'Orazio

文献摘要

被引文献

相似文献

政治和社会科学家监测,分析和预测政治动荡和暴力,防止(或减轻)伤害,并促进全球冲突的管理。他们使用事件编码器系统来实现这一点,该系统从新闻文章中提取结构化表示来设计预测模型和事件驱动的连续监测系统。现有的方法依赖于昂贵的手动注释字典,并且不支持多语言设置。为了推进全球冲突管理,我们提出了一种新的模型,多CoPED(多语言多任务学习BERT编码政治事件数据),通过利用多任务学习和最先进的语言模型编码多语言政治事件。这通过迁移学习利用BERT模型的上下文知识消除了对昂贵字典的需求。多语言实验表明,多CoPED优于现有的事件编码器,提高了23.3%和30.7%的绝对宏观平均F1分数编码事件在英语和西班牙语语料库,分别。我们相信,这种表达能力的提高可以帮助减少对有暴力风险的人的伤害。
Political and social scientists monitor, analyze and predict political unrest and violence, preventing (or mitigating) harm, and promoting the management of global conflict. They do so using event coder systems, which extract structured representations from news articles to design forecast models and event-driven continuous monitoring systems. Existing methods rely on expensive manual annotated dictionaries and do not support multilingual settings. To advance the global conflict management, we propose a novel model, Multi-CoPED (Multilingual Multi-Task Learning BERT for Coding Political Event Data), by exploiting multi-task learning and state-of-the-art language models for coding multilingual political events. This eliminates the need for expensive dictionaries by leveraging BERT models' contextual knowledge through transfer learning. The multilingual experiments demonstrate the superiority of Multi-CoPED over existing event coders, improving the absolute macro-averaged F1-scores by 23.3% and 30.7% for coding events in English and Spanish corpus, respectively. We believe that such expressive performance improvements can help to reduce harms to people at risk of violence.