TextShield: Robust Text Classification Based on Multimodal Embedding and Neural Machine Translation

TextShield: Robust Text Classification Based on Multimodal Embedding and Neural Machine Translation
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Jinfeng Li;Tianyu Du;S. Ji;Rong Zhang;Quan Lu;Min Yang;Ting Wang
Jinfeng Li;Tianyu Du;S. Ji;Rong Zhang;Quan Lu;Min Yang;Ting Wang
中科院分区:
其他
文献类型:
--
作者:
Jinfeng Li;Tianyu Du;S. Ji;Rong Zhang;Quan Lu;Min Yang;Ting Wang

文献摘要

被引文献

相似文献

基于文本的有毒内容检测是减少在线社交媒体环境中有害交互的重要工具。然而,其底层机制,基于深度学习的文本分类(DLTC),本质上容易受到恶意制作的对抗性文本的攻击。为了减轻这种脆弱性,对加强基于英语的DLTC模型进行了深入研究。然而,现有的防御是不有效的基于中文的DLTC模型,由于独特的稀疏性,多样性和变化的中文语言。在本文中,我们通过介绍T EXT S HIELD来弥合这一惊人的差距,这是一种专门为基于中文的DLTC模型设计的新的对抗性防御框架。文本S HIELD在几个关键方面与以前的工作不同:(i)通用-它适用于任何基于中文的DLTC模型,而无需重新训练;(ii)鲁棒性-即使在自适应攻击的设置下,它也显着降低了攻击成功率;(iii)准确-它对合法输入的DLTC模型的性能几乎没有影响。广泛的评估表明,它优于现有的方法和行业领先的平台。今后的工作将探讨其在更广泛的实际任务中的适用性。
Text-based toxic content detection is an important tool for reducing harmful interactions in online social media environments. Yet, its underlying mechanism, deep learning-based text classification (DLTC), is inherently vulnerable to maliciously crafted adversarial texts. To mitigate such vulnerabilities, intensive research has been conducted on strengthening English-based DLTC models. However, the existing defenses are not effective for Chinese-based DLTC models, due to the unique sparseness, diversity, and variation of the Chinese language. In this paper, we bridge this striking gap by presenting T EXT S HIELD , a new adversarial defense framework specifically designed for Chinese-based DLTC models. T EXT S HIELD differs from previous work in several key aspects: (i) generic – it applies to any Chinese-based DLTC models without requiring re-training; (ii) robust – it signifi-cantly reduces the attack success rate even under the setting of adaptive attacks; and (iii) accurate – it has little impact on the performance of DLTC models over legitimate inputs. Extensive evaluations show that it outperforms both existing methods and the industry-leading platforms. Future work will explore its applicability in broader practical tasks.