Integrating Rankings into Quantized Scores in Peer Review

Integrating Rankings into Quantized Scores in Peer Review
复制标题

DOI:
10.48550/arxiv.2204.03505
复制
发表时间:
2022-04
期刊:
Trans. Mach. Learn. Res.
影响因子:
--
通讯作者:
Yusha Liu;Yichong Xu;Nihar B. Shah;Aarti Singh
Yusha Liu;Yichong Xu;Nihar B. Shah;Aarti Singh
中科院分区:
其他
文献类型:
--
作者:
Yusha Liu;Yichong Xu;Nihar B. Shah;Aarti Singh

文献摘要

相似文献

在同行评审中,审稿人通常被要求提供论文的分数。然后,地区主席或项目主席在决策过程中以各种方式使用这些分数。分数通常以量化的形式得出,以适应人类有限的认知能力,用数值来描述他们的观点。研究发现,量化后的分数存在大量的关联度,从而导致大量的信息丢失。为了缓解这个问题,会议已经开始要求审稿人额外提供他们审稿的论文排名。然而,有两个关键的挑战。首先,没有使用排名信息的标准程序,领域主席可能会以不同的方式使用它(包括简单地忽略它们),从而导致同行评审过程中的随意性。其次,没有合适的接口来明智地使用这些数据,也没有方法将其合并到现有的工作流中,从而导致效率低下。我们采用原则性的方法将排名信息整合到分数中。我们方法的输出是与每个评论相关的更新分数,其中也包含了排名。我们的方法解决了上述两个挑战:(i)确保所有论文的排名都以相同的方式纳入更新分数,从而减少随意性;(ii)允许无缝使用为分数设计的现有界面和工作流程。我们在合成数据集以及ICLR 2017会议的同行评议上对我们的方法进行了实证评估,发现与ICLR 2017数据上表现最好的基线相比,它将误差减少了大约30%。
In peer review, reviewers are usually asked to provide scores for the papers. The scores are then used by Area Chairs or Program Chairs in various ways in the decision-making process. The scores are usually elicited in a quantized form to accommodate the limited cognitive ability of humans to describe their opinions in numerical values. It has been found that the quantized scores suffer from a large number of ties, thereby leading to a significant loss of information. To mitigate this issue, conferences have started to ask reviewers to additionally provide a ranking of the papers they have reviewed. There are however two key challenges. First, there is no standard procedure for using this ranking information and Area Chairs may use it in different ways (including simply ignoring them), thereby leading to arbitrariness in the peer-review process. Second, there are no suitable interfaces for judicious use of this data nor methods to incorporate it in existing workflows, thereby leading to inefficiencies. We take a principled approach to integrate the ranking information into the scores. The output of our method is an updated score pertaining to each review that also incorporates the rankings. Our approach addresses the two aforementioned challenges by: (i) ensuring that rankings are incorporated into the updates scores in the same manner for all papers, thereby mitigating arbitrariness, and (ii) allowing to seamlessly use existing interfaces and workflows designed for scores. We empirically evaluate our method on synthetic datasets as well as on peer reviews from the ICLR 2017 conference, and find that it reduces the error by approximately 30% as compared to the best performing baseline on the ICLR 2017 data.