Common Flaws in Running Human Evaluation Experiments in NLP

Common Flaws in Running Human Evaluation Experiments in NLP
复制标题

NLP 人类评估实验的常见缺陷

DOI:
10.1162/coli_a_00508
复制
发表时间:
2024
影响因子:
9.3
通讯作者:
Thomson C
Thomson C
中科院分区:
计算机科学3区
文献类型:
--
作者:
Thomson C

文献摘要

被引文献

相似文献

在进行一组协调的重复人体评估实验时, 在NLP中,我们发现了我们选择纳入的每个实验中的缺陷 通过一个系统的过程。在这个哑炮,我们描述的缺陷类型,我们 发现,其中包括编码错误(例如,加载错误的系统输出 评估),未能遵循标准科学实践(例如,特设 排除参与者和响应),以及报告的数字错误 结果(例如,报告的数字与实验数据不匹配)。如果这些 问题是普遍存在的,这将对严格的 目前正在进行的NLP评估实验。我们讨论研究人员 可以采取哪些措施来减少此类缺陷的发生,包括预注册, 更好的代码开发实践,增加测试和试点,以及 出版后纠正错误。
While conducting a coordinated set of repeat runs of human evaluation experiments in NLP, we discovered flaws in every single experiment we selected for inclusion via a systematic process. In this squib, we describe the types of flaws we discovered, which include coding errors (e.g., loading the wrong system outputs to evaluate), failure to follow standard scientific practice (e.g., ad hoc exclusion of participants and responses), and mistakes in reported numerical results (e.g., reported numbers not matching experimental data). If these problems are widespread, it would have worrying implications for the rigor of NLP evaluation experiments as currently conducted. We discuss what researchers can do to reduce the occurrence of such flaws, including pre-registration, better code development practices, increased testing and piloting, and post-publication addressing of errors.