Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?
复制标题

DOI:
10.18653/v1/2021.acl-long.346
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Pedro Rodriguez;Joe Barrow;Alexander Miserlis Hoyle;John P. Lalor;Robin Jia;Jordan L. Boyd-Graber
Pedro Rodriguez;Joe Barrow;Alexander Miserlis Hoyle;John P. Lalor;Robin Jia;Jordan L. Boyd-Graber
中科院分区:
其他
文献类型:
--
作者:
Pedro Rodriguez;Joe Barrow;Alexander Miserlis Hoyle;John P. Lalor;Robin Jia;Jordan L. Boyd-Graber

文献摘要

被引文献

相似文献

排行榜在自然语言处理领域得到了广泛的应用,推动了该领域的发展。虽然排行榜是对NLP模型的直接排名,但这种简单性可以掩盖评估项目(示例)和主题(NLP模型)的细微差别。与其取代排行榜,我们主张重新想象,以便它们更好地突出是否以及在哪里取得了进展。在教育测试的基础上,我们创建了一个贝叶斯排行榜模型,其中潜在的学科技能和潜在的项目难度预测正确的回答。利用该模型对排行榜的排名可靠性进行了分析。然后,我们展示了该模型可以指导要注释的内容,识别注释错误,检测过度匹配,并识别信息丰富的示例。最后,我们对未来的基准任务提出建议。
Leaderboards are widely used in NLP and push the field forward. While leaderboards are a straightforward ranking of NLP models, this simplicity can mask nuances in evaluation items (examples) and subjects (NLP models). Rather than replace leaderboards, we advocate a re-imagining so that they better highlight if and where progress is made. Building on educational testing, we create a Bayesian leaderboard model where latent subject skill and latent item difficulty predict correct responses. Using this model, we analyze the ranking reliability of leaderboards. Afterwards, we show the model can guide what to annotate, identify annotation errors, detect overfitting, and identify informative examples. We conclude with recommendations for future benchmark tasks.