Mitigating Covertly Unsafe Text within Natural Language Systems

Mitigating Covertly Unsafe Text within Natural Language Systems
复制标题

DOI:
10.48550/arxiv.2210.09306
复制
发表时间:
2022-10
期刊:
--
影响因子:
--
通讯作者:
Alex Mei;Anisha Kabir;Sharon Levy;Melanie Subbiah;Emily Allaway;J. Judge;D. Patton;Bruce Bimber;K. McKeown;William Yang Wang
Alex Mei;Anisha Kabir;Sharon Levy;Melanie Subbiah;Emily Allaway;J. Judge;D. Patton;Bruce Bimber;K. McKeown;William Yang Wang
中科院分区:
其他
文献类型:
--
作者:
Alex Mei;Anisha Kabir;Sharon Levy;Melanie Subbiah;Emily Allaway;J. Judge;D. Patton;Bruce Bimber;K. McKeown;William Yang Wang

文献摘要

相似文献

智能技术日益普遍的问题是文本安全,因为不受控制的系统可能会向用户生成建议,从而导致受伤或危及生命的后果。然而,生成的可能导致人身伤害的语句的明确程度各不相同。在本文中,我们区分了可能导致人身伤害的文本类型,并建立了一个特别未被充分探索的类别:隐蔽的不安全文本。然后,我们根据系统信息进一步细分此类别,并讨论减少每个子类别中文本生成的解决方案。最终,我们的工作定义了导致人身伤害的隐蔽不安全语言问题,并认为利益相关者和监管机构需要优先考虑这一微妙但危险的问题。我们重点介绍缓解策略,以激励未来的研究人员解决这一具有挑战性的问题,并帮助提高智能系统的安全性。
An increasingly prevalent problem for intelligent technologies is text safety, as uncontrolled systems may generate recommendations to their users that lead to injury or life-threatening consequences. However, the degree of explicitness of a generated statement that can cause physical harm varies. In this paper, we distinguish types of text that can lead to physical harm and establish one particularly underexplored category: covertly unsafe text. Then, we further break down this category with respect to the system's information and discuss solutions to mitigate the generation of text in each of these subcategories. Ultimately, our work defines the problem of covertly unsafe language that causes physical harm and argues that this subtle yet dangerous issue needs to be prioritized by stakeholders and regulators. We highlight mitigation strategies to inspire future researchers to tackle this challenging problem and help improve safety within smart systems.