The Human Evaluation Datasheet 1.0: A Template for Recording Details of Human Evaluation Experiments in NLP

The Human Evaluation Datasheet 1.0: A Template for Recording Details of Human Evaluation Experiments in NLP
复制标题

人类评估数据表 1.0:用于记录 NLP 中人类评估实验细节的模板

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
Anya Belz
Anya Belz
中科院分区:
--
文献类型:
--
作者:
Anastasia Shimorina;Anya Belz

文献摘要

参考文献

被引文献

相似文献

本文介绍了人类评估数据表,一个用于记录自然语言处理(NLP)中个人评估实验细节的模板。人体评价数据表最初从Bender和Friedman(2018)、Mitchell等人(2019)和Gebru等人(2020)的开创性论文中获得灵感,旨在促进充分详细记录人体评价的特性,并进行充分的标准化,以支持可比性、荟萃评价和再现性测试。
This paper introduces the Human Evaluation Datasheet, a template for recording the details of individual human evaluation experiments in Natural Language Processing (NLP). Originally taking inspiration from seminal papers by Bender and Friedman (2018), Mitchell et al. (2019), and Gebru et al. (2020), the Human Evaluation Datasheet is intended to facilitate the recording of properties of human evaluations in sufficient detail, and with sufficient standardisation, to support comparability, meta-evaluation, and reproducibility tests.
DOI: 10.18653/v1/2020.inlg-1.23
发表时间: 2020
期刊: --
影响因子: --
作者:
David M. Howcroft;Anya Belz;Miruna Clinciu;Dimitra Gkatzia;Sadid A. Hasan;Saad Mahamood;Simon Mille;Emiel van Miltenburg;Sashank Santhanam;Verena Rieser
通讯作者: David M. Howcroft;Anya Belz;Miruna Clinciu;Dimitra Gkatzia;Sadid A. Hasan;Saad Mahamood;Simon Mille;Emiel van Miltenburg;Sashank Santhanam;Verena Rieser