Rethink reporting of evaluation results in AI
Rethink reporting of evaluation results in AI
复制标题
重新思考人工智能评估结果的报告
DOI:
10.1126/science.adf6369
复制
发表时间:
2023
期刊:
影响因子:
56.9
通讯作者:
Mitchell, Melanie
中科院分区:
文献类型:
--
作者:
Burnell, Ryan;Schellaert, Wout;Burden, John;Ullman, Tomer D.;Martinez-Plumed, Fernando;Tenenbaum, Joshua B.;Rutar, Danaja;Cheke, Lucy G.;Sohl-Dickstein, Jascha;Mitchell, Melanie
Artificial intelligence (AI) systems have begun to be deployed in high-stakes contexts, including autonomous driving and medical diagnosis. In contexts such as these, the consequences of system failures can be devastating. It is therefore vital that researchers and policy-makers have a full understanding of the capabilities and weaknesses of AI systems so that they can make informed decisions about where these systems are safe to use and how they might be improved. Unfortunately, current approaches to AI evaluation make it exceedingly difficult to build such an understanding, for two key reasons. First, aggregate metrics make it hard to predict how a system will perform in a particular situation. Second, the instance-by-instance evaluation results that could be used to unpack these aggregate metrics are rarely made available . Here, we propose a path forward in which results are presented in more nuanced ways and instance-by-instance evaluation results are made publicly available.