Editors’ Introduction to the Special Section on Replicability in Psychological Science

Editors’ Introduction to the Special Section on Replicability in Psychological Science
复制标题

心理科学可重复性专题编辑简介

DOI:
--
复制
发表时间:
2012
影响因子:
12.6
通讯作者:
E. Wagenmakers
E. Wagenmakers
中科院分区:
心理学1区
文献类型:
--
作者:
H. Pashler;E. Wagenmakers

文献摘要

被引文献

相似文献

目前心理科学是否存在信任危机,反映出从业者对该领域研究结果的可靠性产生了前所未有的怀疑?看起来肯定是有的。这些疑虑随着 2011 年一系列令人不快的事件的展开而出现和增长:Diederik Stapel 欺诈案(参见 Stroebe、Postmes 和 Spears,2012 年,本期)、在一家主要社会心理学杂志上发表了一篇文章,旨在展示超感官知觉的证据(Bem,2011 年),随后受到广泛的公众嘲笑(参见 Galak、LeBoeuf、Nelson 和 Simmons,正在出版; Wagenmakers、Wetzels、Borsboom 和 van der Maas,2011 年),Wicherts 及其同事的报告称,心理学家通常不愿意或无法分享他们已发表的数据进行重新分析(Wicherts、Bakker 和 Molenaar,2011 年;另请参阅 Wicherts、Borsboom、Kats 和 Molenaar,2006 年),以及在《心理科学》杂志上发表的一篇重要文章表明了如何容易地分享这些数据。在没有任何实际效果的情况下,研究人员仍然可以通过各种有问题的研究实践(QRP)获得统计显着差异,例如探索多个因变量或协变量,并且仅在产生显着结果时才报告这些结果(Simmons、Nelson 和 Simonsohn,2011)。对于那些预计 2011 年的尴尬很快就会消失在记忆中的心理学家来说,2012 年的情况却迅速恶化,社会认知领域出现了彻头彻尾的欺诈行为的新迹象(Simonsohn,2012),《心理科学》上的一篇文章表明,许多心理学家承认至少参与了 Simmons 及其同事检查的一些 QRP(John、Loewenstein 和 Prelec, 2012),令人不安的新元分析证据表明 Simmons 及其同事描述的 QRP 甚至可能在心理学文献中的 p 值分布中留下明显的迹象(Masicampo & Lalande,出版中;Simonsohn,2012),并且科学杂志和博客中围绕一些研究人员在复制社会认知领域的众所周知的结果时遇到的问题展开了激烈的争论(Bower, 2012;勇,2012)。尽管心理学在这两年期间经历的非常公开的问题让我们这些在该领域工作的人感到尴尬,但一些人感到安慰的是,在同一时期,整个科学界也出现了类似的担忧(由稍后将描述的启示引发)。一些可疑的不可重复性原因,例如发表偏见(只发表正面研究结果的倾向),已经讨论了多年;事实上,文件抽屉问题这个词是几十年前由一位杰出的心理学家首次创造的(Rosenthal,1979)。然而,许多人推测,近年来,随着学术界过度竞争的学术氛围和激励计划,这些问题变得更加严重,该激励计划为过度推销自己的工作提供丰厚的奖励,而对谨慎和谨慎的人则几乎没有奖励(参见 Giner-Sorolla,2012 年,本期)。同样令人不安的是,调查人员重复彼此工作的频率似乎比过去更少,这可能再次反映出激励计划的歪曲(这一点在本期的几篇文章中讨论过,例如 Makel、Plucker 和 Hegarty,2012 年)。目前尚不清楚心理学文献中出现错误的频率,但许多事实表明它可能高得令人不安。 Ioannidis(2005)通过简单的数学模型表明,任何忽视复制的科学领域都很容易陷入悲惨的境地(正如他最著名的文章的标题所说)“大多数发表的研究结果都是错误的”(另见 Ioannidis,2012,本期,以及 Pashler & Harris,2012,本期)。与此同时,癌症研究中出现的报告使这种严峻的情况看起来更加可信:2012年,几家大型制药公司透露,他们为复制已发表的癌症生物学学术研究中令人兴奋的临床前发现所做的努力很少验证原始结果(Begley & Ellis, 2012;另见 Osherovich, 2011;Prinz, Schlange, & Asadullah, 2011)。
Is there currently a crisis of confidence in psychological science reflecting an unprecedented level of doubt among practitioners about the reliability of research findings in the field? It would certainly appear that there is. These doubts emerged and grew as a series of unhappy events unfolded in 2011: the Diederik Stapel fraud case (see Stroebe, Postmes, & Spears, 2012, this issue), the publication in a major social psychology journal of an article purporting to show evidence of extrasensory perception (Bem, 2011) followed by widespread public mockery (see Galak, LeBoeuf, Nelson, & Simmons, in press; Wagenmakers, Wetzels, Borsboom, & van der Maas, 2011), reports by Wicherts and colleagues that psychologists are often unwilling or unable to share their published data for reanalysis (Wicherts, Bakker, & Molenaar, 2011; see also Wicherts, Borsboom, Kats, & Molenaar, 2006), and the publication of an important article in Psychological Science showing how easily researchers can, in the absence of any real effects, nonetheless obtain statistically significant differences through various questionable research practices (QRPs) such as exploring multiple dependent variables or covariates and only reporting these when they yield significant results (Simmons, Nelson, & Simonsohn, 2011). For those psychologists who expected that the embarrassments of 2011 would soon recede into memory, 2012 offered instead a quick plunge from bad to worse, with new indications of outright fraud in the field of social cognition (Simonsohn, 2012), an article in Psychological Science showing that many psychologists admit to engaging in at least some of the QRPs examined by Simmons and colleagues (John, Loewenstein, & Prelec, 2012), troubling new meta-analytic evidence suggesting that the QRPs described by Simmons and colleagues may even be leaving telltale signs visible in the distribution of p values in the psychological literature (Masicampo & Lalande, in press; Simonsohn, 2012), and an acrimonious dust-up in science magazines and blogs centered around the problems some investigators were having in replicating well-known results from the field of social cognition (Bower, 2012; Yong, 2012). Although the very public problems experienced by psychology over this 2-year period are embarrassing to those of us working in the field, some have found comfort in the fact that, over the same period, similar concerns have been arising across the scientific landscape (triggered by revelations that will be described shortly). Some of the suspected causes of unreplicability, such as publication bias (the tendency to publish only positive findings) have been discussed for years; in fact, the phrase file-drawer problem was first coined by a distinguished psychologist several decades ago (Rosenthal, 1979). However, many have speculated that these problems have been exacerbated in recent years as academia reaps the harvest of a hypercompetitive academic climate and an incentive scheme that provides rich rewards for overselling one’s work and few rewards at all for caution and circumspection (see Giner-Sorolla, 2012, this issue). Equally disturbing, investigators seem to be replicating each others’ work even less often than they did in the past, again presumably reflecting an incentive scheme gone askew (a point discussed in several articles in this issue, e.g., Makel, Plucker, & Hegarty, 2012). The frequency with which errors appear in the psychological literature is not presently known, but a number of facts suggest it might be disturbingly high. Ioannidis (2005) has shown through simple mathematical modeling that any scientific field that ignores replication can easily come to the miserable state wherein (as the title of his most famous article puts it) “most published research findings are false” (see also Ioannidis, 2012, this issue, and Pashler & Harris, 2012, this issue). Meanwhile, reports emerging from cancer research have made such grim scenarios seem more plausible: In 2012, several large pharmaceutical companies revealed that their efforts to replicate exciting preclinical findings from published academic studies in cancer biology were only rarely verifying the original results (Begley & Ellis, 2012; see also Osherovich, 2011; Prinz, Schlange, & Asadullah, 2011).