Rethink reporting of evaluation results in AI

Rethink reporting of evaluation results in AI
复制标题

重新思考人工智能评估结果的报告

DOI:
10.1126/science.adf6369
复制
发表时间:
2023
期刊:
影响因子:
56.9
通讯作者:
Mitchell, Melanie
Mitchell, Melanie
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Burnell, Ryan;Schellaert, Wout;Burden, John;Ullman, Tomer D.;Martinez-Plumed, Fernando;Tenenbaum, Joshua B.;Rutar, Danaja;Cheke, Lucy G.;Sohl-Dickstein, Jascha;Mitchell, Melanie

文献摘要

被引文献

相似文献

人工智能(AI)系统已经开始部署在高风险的环境中,包括自动驾驶和医疗诊断。在这样的情况下,系统故障的后果可能是毁灭性的。因此,至关重要的是,研究人员和政策制定者充分了解人工智能系统的能力和弱点,以便他们能够就这些系统在哪里安全使用以及如何改进做出明智的决定。不幸的是,目前的人工智能评估方法使建立这样的理解变得极其困难,原因有两个。首先,聚合指标使得很难预测系统在特定情况下将如何运行。其次,可以用来解包这些聚合指标的逐个实例的评估结果很少提供。在这里,我们提出了一条前进的道路,其中以更细微的方式呈现结果,并将逐个实例的评估结果公之于众。
Artificial intelligence (AI) systems have begun to be deployed in high-stakes contexts, including autonomous driving and medical diagnosis. In contexts such as these, the consequences of system failures can be devastating. It is therefore vital that researchers and policy-makers have a full understanding of the capabilities and weaknesses of AI systems so that they can make informed decisions about where these systems are safe to use and how they might be improved. Unfortunately, current approaches to AI evaluation make it exceedingly difficult to build such an understanding, for two key reasons. First, aggregate metrics make it hard to predict how a system will perform in a particular situation. Second, the instance-by-instance evaluation results that could be used to unpack these aggregate metrics are rarely made available . Here, we propose a path forward in which results are presented in more nuanced ways and instance-by-instance evaluation results are made publicly available.