Deep Headline Generation for Clickbait Detection

Deep Headline Generation for Clickbait Detection
复制标题

DOI:
10.1109/icdm.2018.00062
复制
发表时间:
2018-11
期刊:
2018 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Kai Shu;Suhang Wang;Thai Le;Dongwon Lee;Huan Liu
Kai Shu;Suhang Wang;Thai Le;Dongwon Lee;Huan Liu
中科院分区:
其他
文献类型:
--
作者:
Kai Shu;Suhang Wang;Thai Le;Dongwon Lee;Huan Liu

文献摘要

被引文献

相似文献

点击诱饵是吸引读者点击的社交帖子或耸人听闻的标题。点击诱饵在社交媒体上普遍存在,可能对用户和媒体生态系统产生重大负面影响。例如,用户可能会被误导而接收不准确的信息或陷入点击劫持攻击。同样,由于点击诱饵的盛行,媒体平台可能会失去读者的信任和收入。要使用监督学习框架通过计算检测社交媒体上的此类点击诱饵,主要障碍之一是由于标记成本高昂而缺乏大规模标记训练数据。随着深度生成模型的最新进展,为了应对这一挑战,我们建议生成具有特定样式的合成标题,并探索其实用程序以帮助改进点击诱饵检测。特别是,我们建议通过风格迁移从原始文档生成风格化的标题。此外,由于文本的离散性以及在实现风格迁移的同时保留文档语义的要求等几个挑战,生成风格化标题并非易事,因此我们提出了一种新颖的解决方案,称为风格化标题生成(SHG),它不仅可以生成可读且真实的标题来扩大原始训练数据,而且有助于提高监督学习的分类能力。真实数据集上的实验结果证明了 SHG 在生成用于点击诱饵检测的高质量和高实用性标题方面的有效性。
Clickbaits are catchy social posts or sensational headlines that attempt to lure readers to click. Clickbaits are pervasive on social media and can have significant negative impacts on both users and media ecosystems. For example, users may be misled to receive inaccurate information or fall into click-jacking attacks. Similarly, media platforms could lose readers' trust and revenues due to the prevalence of clickbaits. To computationally detect such clickbaits on social media using a supervised learning framework, one of the major obstacles is the lack of large-scale labeled training data, due to the high cost of labeling. With the recent advancements of deep generative models, to address this challenge, we propose to generate synthetic headlines with specific styles and explore their utilities to help improve clickbait detection. In particular, we propose to generate stylized headlines from original documents with style transfer. Furthermore, as it is non-trivial to generate stylized headlines due to several challenges such as the discrete nature of texts and the requirements of preserving semantic meaning of document while achieving style transfer, we propose a novel solution, named as Stylized Headline Generation (SHG), that can not only generate readable and realistic headlines to enlarge original training data, but also help improve the classification capacity of supervised learning. The experimental results on real-world datasets demonstrate the effectiveness of SHG in generating high-quality and high-utility headlines for clickbait detection.