Detecting Trending Terms in Cybersecurity Forum Discussions

Detecting Trending Terms in Cybersecurity Forum Discussions
复制标题

DOI:
10.18653/v1/2020.wnut-1.15
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Jack Hughes;S. Aycock;Andrew Caines;P. Buttery;Alice Hutchings
Jack Hughes;S. Aycock;Andrew Caines;P. Buttery;Alice Hutchings
中科院分区:
其他
文献类型:
--
作者:
Jack Hughes;S. Aycock;Andrew Caines;P. Buttery;Alice Hutchings

文献摘要

相似文献

我们提出了一种轻量级方法,用于识别与已知术语先验相关的当前趋势术语,使用具有信息先验的加权对数赔率比。我们将这种方法应用于一个英文地下黑客论坛的帖子数据集,这些帖子跨越了十多年的活动,其中的帖子包含拼写错误、拼写变体、首字母缩写和俚语。我们的统计方法支持分析语言变化和讨论主题随时间的推移,而不需要为每个时间间隔训练主题模型进行分析。我们通过将结果与使用带有人类标注的折扣累积增益度量的TF-IDF进行比较来评估该方法,发现我们的方法在信息检索方面优于TF-IDF。
We present a lightweight method for identifying currently trending terms in relation to a known prior of terms, using a weighted log-odds ratio with an informative prior. We apply this method to a dataset of posts from an English-language underground hacking forum, spanning over ten years of activity, with posts containing misspellings, orthographic variation, acronyms, and slang. Our statistical approach supports analysis of linguistic change and discussion topics over time, without a requirement to train a topic model for each time interval for analysis. We evaluate the approach by comparing the results to TF-IDF using the discounted cumulative gain metric with human annotations, finding our method outperforms TF-IDF on information retrieval.