On Testing for Biases in Peer Review

On Testing for Biases in Peer Review
复制标题

DOI:
--
复制
发表时间:
2019-12
期刊:
--
影响因子:
--
通讯作者:
Ivan Stelmakh;Nihar B. Shah;Aarti Singh
Ivan Stelmakh;Nihar B. Shah;Aarti Singh
中科院分区:
其他
文献类型:
--
作者:
Ivan Stelmakh;Nihar B. Shah;Aarti Singh

文献摘要

被引文献

相似文献

我们考虑学术研究中的偏见问题,特别是在同行评议中。将作者身份暴露给评论家是否会导致对某些群体的偏见,这是一个长期存在的争论,我们的重点是设计测试来检测这种偏见的存在。我们的出发点是Tomkins、Zhang和Heavlin最近的一项引人注目的工作,他们进行了一项受控的大规模实验,以调查WSDM会议同行审查中存在的偏见。在本文中我们给出了两组结果。第一组结果是否定的,与Tomkins等人工作中使用的统计测试和实验设置有关。我们表明,所使用的检验不能保证对虚警概率的控制,并且在相关变量之间的相关性下,再加上以下任何一种高概率的条件,可以在实际上不存在偏差的情况下宣告存在偏差:(A)测量误差,(B)模型失配,(C)评审者校准。此外,我们还表明,如果(D)竞标是以非盲目的方式进行的,或者(E)采用流行的评审员分配程序,则他们的实验设置本身可能会夸大错误警报概率。我们的第二组结果是积极的,因为我们提出了一个在(单盲与双盲)同行审查中测试偏差的一般框架。然后,我们给出了一个假设检验,即使在(A)-(C)条件下,也能保证对虚警概率和非平凡功率的控制。条件(D)和(E)是更基本的问题,与实验设置有关,与测试无关。
We consider the issue of biases in scholarly research, specifically, in peer review. There is a long standing debate on whether exposing author identities to reviewers induces biases against certain groups, and our focus is on designing tests to detect the presence of such biases. Our starting point is a remarkable recent work by Tomkins, Zhang and Heavlin which conducted a controlled, large-scale experiment to investigate existence of biases in the peer reviewing of the WSDM conference. We present two sets of results in this paper. The first set of results is negative, and pertains to the statistical tests and the experimental setup used in the work of Tomkins et al. We show that the test employed therein does not guarantee control over false alarm probability and under correlations between relevant variables, coupled with any of the following conditions, with high probability can declare a presence of bias when it is in fact absent: (a) measurement error, (b) model mismatch, (c) reviewer calibration. Moreover, we show that the setup of their experiment may itself inflate false alarm probability if (d) bidding is performed in non-blind manner or (e) popular reviewer assignment procedure is employed. Our second set of results is positive, in that we present a general framework for testing for biases in (single vs. double blind) peer review. We then present a hypothesis test with guaranteed control over false alarm probability and non-trivial power even under conditions (a)--(c). Conditions (d) and (e) are more fundamental problems that are tied to the experimental setup and not necessarily related to the test.