Codewords Detection in Microblogs Focusing on Differences in Word Use Between Two Corpora

Codewords Detection in Microblogs Focusing on Differences in Word Use Between Two Corpora
复制标题

DOI:
10.1109/iccece49321.2020.9231109
复制
发表时间:
2020-08
期刊:
2020 International Conference on Computing, Electronics & Communications Engineering (iCCECE)
影响因子:
--
通讯作者:
Takuro Hada;Y. Sei;Yasuyuki Tahara;Akihiko Ohsuga
Takuro Hada;Y. Sei;Yasuyuki Tahara;Akihiko Ohsuga
中科院分区:
其他
文献类型:
--
作者:
Takuro Hada;Y. Sei;Yasuyuki Tahara;Akihiko Ohsuga

文献摘要

相似文献

近年来,利用微博进行毒品交易的现象有所上升,并成为一个社会问题。打击贩毒等犯罪的网络巡逻的一种常见方法是搜索与犯罪相关的关键字。然而,发布犯罪诱导信息的犯罪分子最大限度地使用“码字”而不是关键词,如enjo kosai,大麻和甲基苯丙胺,以掩饰他们的犯罪意图。研究表明,这些码字一旦变得流行就会改变;因此,搜索特定单词需要花费大量精力来跟踪最新的码字。在这项研究中,我们专注于码字的外观和那些可能被包括在犯罪职位,旨在检测码字的高可能性包括在犯罪职位。我们提出了新的方法来检测码字的基础上的差异,字的使用和隐藏字检测进行实验,以评估方法的有效性。实验结果表明,该方法能够检测出初始列表以外的隐藏词,并且相对于基线方法有更好的检测效果。这些发现证明了所提出的方法能够快速、自动地检测随时间变化的码字和诱发犯罪的博客文章,从而有可能减轻对码字进行持续监控的负担。
In recent years, drug trafficking using microblogs has risen and become a social problem. A common method of cyber patrols for cracking down on crimes, such as drug trafficking, involves searching for crime-related keywords. However, criminals who post crime-inducing messages make maximum use of "codewords" rather than keywords, such as enjo kosai, marijuana, and methamphetamine, to camouflage their criminal intentions. Research suggests that these codewords change once they become popular; therefore, searching for a specific word requires significant effort to keep track of the latest codewords. In this study, we focused on the appearance of codewords and those likely to be included in incriminating posts with aim to detect codewords with the high likelihood of inclusion in incriminating posts. We proposed new methods for detecting codewords based on differences in word usage and conducted experiments on concealed-word detection in order to evaluate method effectiveness. The results showed that the proposed method was capable of detecting concealed words other than those in the initial list and to better degree relative to baseline methods. These findings demonstrated the ability of the proposed method to rapidly and automatically detect codewords that change over time and blog posts that induce crimes, thereby potentially reducing the burden of continuous monitoring of codewords.