Supporting Human-AI Collaboration in Auditing LLMs with LLMs

Supporting Human-AI Collaboration in Auditing LLMs with LLMs
复制标题

DOI:
10.1145/3600211.3604712
复制
发表时间:
2023-04
期刊:
Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society
影响因子:
--
通讯作者:
Charvi Rastogi;Marco Tulio Ribeiro;Nicholas King;Saleema Amershi
Charvi Rastogi;Marco Tulio Ribeiro;Nicholas King;Saleema Amershi
中科院分区:
其他
文献类型:
--
作者:
Charvi Rastogi;Marco Tulio Ribeiro;Nicholas King;Saleema Amershi

文献摘要

被引文献

相似文献

通过在社会技术系统中的部署,大型语言模型 (LLM) 正变得越来越强大和普遍。然而,这些语言模型,无论是分类还是生成,都已被证明是有偏见的,行为不负责任,对人们造成了大规模的伤害。在部署之前严格审核这些语言模型至关重要。现有的审计工具使用人类和人工智能之一或两者来发现故障。在这项工作中,我们借鉴了人类与人工智能协作和意义建构方面的文献,并采访了安全和公平人工智能领域的研究专家,以构建审计工具:AdaTest [36],该工具由生成式法学硕士提供支持。通过设计过程,我们强调了意义建构和人机交互的重要性,以在协作审计中利用人类和生成模型的互补优势。为了评估增强工具 AdaTest++ 的有效性,我们与参与者进行了用户研究,审核了两种商业语言模型:OpenAI 的 GPT-3 和 Azure 的情感分析模型。定性分析表明,AdaTest++ 有效地利用了人类的优势,例如图式化、假设检验。此外,使用我们的工具,用户识别了各种故障模式,涵盖 2 个任务的 26 个不同主题,这些模式已在正式审计中显示,也包括以前未充分报告的故障模式。
Large language models (LLMs) are increasingly becoming all-powerful and pervasive via deployment in sociotechnical systems. Yet these language models, be it for classification or generation, have been shown to be biased, behave irresponsibly, causing harm to people at scale. It is crucial to audit these language models rigorously before deployment. Existing auditing tools use either or both humans and AI to find failures. In this work, we draw upon literature in human-AI collaboration and sensemaking, and interview research experts in safe and fair AI, to build upon the auditing tool: AdaTest [36], which is powered by a generative LLM. Through the design process we highlight the importance of sensemaking and human-AI communication to leverage complementary strengths of humans and generative models in collaborative auditing. To evaluate the effectiveness of AdaTest++, the augmented tool, we conduct user studies with participants auditing two commercial language models: OpenAI’s GPT-3 and Azure’s sentiment analysis model. Qualitative analysis shows that AdaTest++ effectively leverages human strengths such as schematization, hypothesis testing. Further, with our tool, users identified a variety of failures modes, covering 26 different topics over 2 tasks, that have been shown in formal audits and also those previously under-reported.