Defending Against Neural Fake News

Defending Against Neural Fake News
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Rowan Zellers;Ari Holtzman;Hannah Rashkin;Yonatan Bisk;Ali Farhadi;Franziska Roesner;Yejin Choi
Rowan Zellers;Ari Holtzman;Hannah Rashkin;Yonatan Bisk;Ali Farhadi;Franziska Roesner;Yejin Choi
中科院分区:
其他
文献类型:
--
作者:
Rowan Zellers;Ari Holtzman;Hannah Rashkin;Yonatan Bisk;Ali Farhadi;Franziska Roesner;Yejin Choi

文献摘要

被引文献

相似文献

自然语言产生的最新进展引起了双重用途的关注。尽管诸如摘要和翻译之类的应用是积极的,但基础技术也可能使对手能够产生神经假新闻:有针对性的宣传,这些宣传紧密模仿了真实新闻的风格。现代计算机安全依赖于仔细的威胁建模:从对手的角度来确定潜在的威胁和脆弱性,并探索对这些威胁的潜在缓解。同样,开发针对神经假新闻的强大防御能力要求我们首先仔细调查并表征这些模型的风险。因此,我们提出了一个称为Grover的可控文本生成的模型。鉴于诸如“疫苗和自闭症之间的链接”之类的标题,格罗弗可以生成本文的其余部分。人类认为这些世代比人写的虚假信息更值得信赖。开发针对Grover等发电机的强大验证技术至关重要。我们发现,最佳当前歧视者可以通过访问适中的培训数据,从真实的,人文编写的新闻中对神经假新闻进行分类。违反直觉,对格罗弗的最佳防守本身就是格罗弗本身,精度为92%,这表明了公众释放强发电机的重要性。我们进一步研究了这些结果,表明暴露偏见 - 以及减轻其影响的取样策略 - 都留下了类似歧视者可以接受的伪像。最后,我们通过讨论有关该技术的道德问题,并计划公开发布Grover,为更好地检测神经假新闻铺平道路。
Recent progress in natural language generation has raised dual-use concerns. While applications like summarization and translation are positive, the underlying technology also might enable adversaries to generate neural fake news: targeted propaganda that closely mimics the style of real news. Modern computer security relies on careful threat modeling: identifying potential threats and vulnerabilities from an adversary's point of view, and exploring potential mitigations to these threats. Likewise, developing robust defenses against neural fake news requires us first to carefully investigate and characterize the risks of these models. We thus present a model for controllable text generation called Grover. Given a headline like `Link Found Between Vaccines and Autism,' Grover can generate the rest of the article; humans find these generations to be more trustworthy than human-written disinformation. Developing robust verification techniques against generators like Grover is critical. We find that best current discriminators can classify neural fake news from real, human-written, news with 73% accuracy, assuming access to a moderate level of training data. Counterintuitively, the best defense against Grover turns out to be Grover itself, with 92% accuracy, demonstrating the importance of public release of strong generators. We investigate these results further, showing that exposure bias -- and sampling strategies that alleviate its effects -- both leave artifacts that similar discriminators can pick up on. We conclude by discussing ethical issues regarding the technology, and plan to release Grover publicly, helping pave the way for better detection of neural fake news.