Context-specific Language Modeling for Human Trafficking Detection from Online Advertisements

Context-specific Language Modeling for Human Trafficking Detection from Online Advertisements
复制标题

用于在线广告人口贩运检测的上下文特定语言模型

DOI:
10.18653/v1/p19-1114
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Andy E. Fano
Andy E. Fano
中科院分区:
--
文献类型:
--
作者:
Saeideh Shahrokh Esfahani;Michael J. Cafarella;M. Pouyan;Gregory J. DeAngelo;E. Eneva;Andy E. Fano

文献摘要

被引文献

相似文献

人口贩运是一个全球性的危机。人贩子通过在线广告匿名提供性服务来剥削受害者。这些广告通常包含一些线索,执法部门可以利用这些线索将潜在的人口贩卖案件与自愿性广告区分开来。问题在于,广告的数量对于人工处理来说太大了。理想情况下,可以使用集中式半自动工具来协助执法机构完成这项任务。在这里,我们提出了一种使用自然语言处理来识别这些网站上的贩运广告的方法。我们提出了一个通过集成多个文本特征集的分类器,包括公开可用的预训练文本语言模型双向编码器表示从变压器(BERT)。在本文中,我们证明了使用该组合特征集的分类器比单独使用任何单个特征集具有更好的性能。
Human trafficking is a worldwide crisis. Traffickers exploit their victims by anonymously offering sexual services through online advertisements. These ads often contain clues that law enforcement can use to separate out potential trafficking cases from volunteer sex advertisements. The problem is that the sheer volume of ads is too overwhelming for manual processing. Ideally, a centralized semi-automated tool can be used to assist law enforcement agencies with this task. Here, we present an approach using natural language processing to identify trafficking ads on these websites. We propose a classifier by integrating multiple text feature sets, including the publicly available pre-trained textual language model Bi-directional Encoder Representation from transformers (BERT). In this paper, we demonstrate that a classifier using this composite feature set has significantly better performance compared to any single feature set alone.