Do they agree? Bibliometric evaluation versus informed peer review in the Italian research assessment exercise

Do they agree? Bibliometric evaluation versus informed peer review in the Italian research assessment exercise
复制标题

DOI:
10.1007/s11192-016-1929-y
复制
发表时间:
2016-09-01
期刊:
影响因子:
3.9
通讯作者:
De Nicolao, Giuseppe
De Nicolao, Giuseppe
中科院分区:
管理学3区
文献类型:
--
作者:
Baccini, Alberto;De Nicolao, Giuseppe

文献摘要

被引文献

相似文献

在意大利的研究评估工作中,国家机构ANVUR进行了一项实验,以评估知情同行评议(IR)和文献计量学归因于期刊论文的等级之间的一致性。用这两种方法对样本进行评价,并用加权Cohen‘s Kappas法对一致性进行分析。ANVUR提出的结果表明总体上是“良好的”或“足够的”同意。本文根据现有的解释卡伯值的统计准则重新检验了实验结果,结果表明,对于所有研究领域来说,一致性程度(总是在0.09-0.42的范围内)必须被解释为不可接受、较差或在少数情况下至多是公平的。唯一值得注意的例外情况也得到了统计荟萃分析的证实,那就是经济学和统计学(领域13)及其子领域的适度一致。我们表明,在第13区通过的实验方案相对于所有其他研究领域都有很大的修改,以至于经济学和统计学的结果必须被认为是致命的缺陷。一致性差的证据支持这样的结论,即IR和文献计量学不会产生类似的结果,在意大利的研究评估中采用这两种方法可能会在其最终结果中引入系统性和未知的偏差。ANVUR得出的结论必须颠倒:现有证据根本不能证明在同一项研究评估工作中联合使用IR和文献计量学是合理的。
During the Italian research assessment exercise, the national agency ANVUR performed an experiment to assess agreement between grades attributed to journal articles by informed peer review (IR) and by bibliometrics. A sample of articles was evaluated by using both methods and agreement was analyzed by weighted Cohen's kappas. ANVUR presented results as indicating an overall "good'' or "more than adequate'' agreement. This paper re-examines the experiment results according to the available statistical guidelines for interpreting kappa values, by showing that the degree of agreement (always in the range 0.09-0.42) has to be interpreted, for all research fields, as unacceptable, poor or, in a few cases, as, at most, fair. The only notable exception, confirmed also by a statistical meta-analysis, was a moderate agreement for economics and statistics (Area 13) and its sub-fields. We show that the experiment protocol adopted in Area 13 was substantially modified with respect to all the other research fields, to the point that results for economics and statistics have to be considered as fatally flawed. The evidence of a poor agreement supports the conclusion that IR and bibliometrics do not produce similar results, and that the adoption of both methods in the Italian research assessment possibly introduced systematic and unknown biases in its final results. The conclusion reached by ANVUR must be reversed: the available evidence does not justify at all the joint use of IR and bibliometrics within the same research assessment exercise.