Towards Dark Jargon Interpretation in Underground Forums

Towards Dark Jargon Interpretation in Underground Forums
复制标题

DOI:
10.1007/978-3-030-72240-1_40
复制
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Dominic Seyler;Wei Liu;Xiaofeng Wang;Chengxiang Zhai
Dominic Seyler;Wei Liu;Xiaofeng Wang;Chengxiang Zhai
中科院分区:
其他
文献类型:
--
作者:
Dominic Seyler;Wei Liu;Xiaofeng Wang;Chengxiang Zhai

文献摘要

相似文献

黑话是一些看起来很好的词,但却有着隐藏的、邪恶的含义,被地下论坛的参与者用于非法行为。例如,黑暗的术语“老鼠”经常被用来代替“远程木马”。在这项工作中,我们提出了一种新的方法自动识别和解释黑暗的行话。我们把这个问题形式化为一个从黑暗的词到没有隐藏意义的“干净”词的映射。我们的方法利用可解释的表示黑暗和干净的话的概率分布的形式在一个共享的词汇。在我们的实验中,我们表明我们的方法是有效的,在黑暗的行话识别,因为它优于另一个基线的模拟数据。使用手动评估,我们表明,我们的方法是能够检测在现实世界的地下论坛数据集的黑暗行话。
Dark jargons are benign-looking words that have hidden, sinister meanings and are used by participants of underground forums for illicit behavior. For example, the dark term “rat” is often used in lieu of “RemoteAccessTrojan”. In this work we present a novel method towards automatically identifying and interpreting dark jargons. We formalize the problem as a mapping from dark words to “clean” words with no hidden meaning. Our method makes use of interpretable representations of dark and clean words in the form of probability distributions over a shared vocabulary. In our experiments we show our method to be effective in terms of dark jargon identification, as it outperforms another baseline on simulated data. Using manual evaluation, we show that our method is able to detect dark jargons in a real-world underground forum dataset.