Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts

Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts
复制标题

DOI:
10.18653/v1/d18-1386
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Samuel Carton;Qiaozhu Mei;P. Resnick
Samuel Carton;Qiaozhu Mei;P. Resnick
中科院分区:
其他
文献类型:
--
作者:
Samuel Carton;Qiaozhu Mei;P. Resnick

文献摘要

被引文献

相似文献

我们引入了一种对抗性方法,用于生成神经文本分类器决策的高召回率解释。在通过硬注意力进行提取解释的现有架构的基础上,我们添加了一个对抗层,用于扫描注意力的剩余部分以获取剩余的预测信号。受检测社交媒体评论中的人身攻击这一重要领域的推动,我们还证明了通过显式操纵其偏差项来为模型手动设置语义上适当的“默认”行为的重要性。我们开发了一组人工注释的人身攻击验证集来评估这些变化的影响。
We introduce an adversarial method for producing high-recall explanations of neural text classifier decisions. Building on an existing architecture for extractive explanations via hard attention, we add an adversarial layer which scans the residual of the attention for remaining predictive signal. Motivated by the important domain of detecting personal attacks in social media comments, we additionally demonstrate the importance of manually setting a semantically appropriate “default” behavior for the model by explicitly manipulating its bias term. We develop a validation set of human-annotated personal attacks to evaluate the impact of these changes.