Disturbed by meta-analysis?

Disturbed by meta-analysis?
复制标题

受到荟萃分析的困扰?

DOI:
--
复制
发表时间:
1988
期刊:
影响因子:
56.9
通讯作者:
K. Wachter
K. Wachter
中科院分区:
综合性期刊1区
文献类型:
--
作者:
K. Wachter

文献摘要

被引文献

相似文献

我!很久以前,《水星》杂志上有一篇名为《最疯狂的事》的文章,对已发表的关于精神疾病表现与月相之间存在关联的说法进行了研究。作者j·罗顿(J. Rotton)和i·凯利(I. Kelly)在不同的月相上发现了不同的研究,这些研究在统计上具有重要意义。“为了解决这个问题,”他们写道,“我们采用了荟萃分析,这是一种综合不同研究结果的统计程序。在我们的荟萃分析中,我们没有发现关于满月影响的普遍信念的证据”(1,第75和95页)。越来越多来自各个领域的科学家在回顾科学文献时都采用“元分析”。Glass(2)在1976年创造了“元分析”这个名字;它指的是研究综合,使用正式的统计程序来检索、选择和组合以前独立研究的结果。荟萃分析的繁荣正在兴起,但这种繁荣并没有受到普遍欢迎。我的一些朋友很乐意把元分析本身归为罗顿和凯利所研究的与月相关系的精神错乱。元分析发展迅速。我们能指望它像无常的月亮一样迅速衰落吗?我认为,我朋友们的怀疑主要集中在四个方面。首先,人们怀疑“垃圾进出”在这里被花哨的统计数据所掩盖。第二种观点认为,如果用为实验结果的统计样本而发明的程序来处理各种各样的研究结果,就会犯一个基本的错误,而没有受控条件、均匀测量量表和使程序有效的统计独立性的好处。第三,人们担心,对以往研究成果的研究正被简化为一项常规的编码任务,交给研究助理来完成,通过压制智慧的作用来提高每个作者每月的产出。与这种恐惧相关的是第四种感觉,即荟萃分析适应了一个坏科学以数字的重量驱逐好科学的世界。第一种指控——伪装指控——显然是错误的。在荟萃分析中几乎没有花哨的统计数据。教科书的主题是非常基本的,脚踏实地,反对一切神秘化。我最喜欢的例子是赫奇斯和奥尔金、莱特和皮勒默、罗森塔尔和沃尔夫的作品。也没有太多的伪装倾向。当一个人试图在一项又一项研究的可比基础上提取或计算影响的大小时,糟糕的研究就会暴露得更明显,而非正式的研究审稿人通常依赖于作者自己对结论的总结。荟萃分析的作者永远在哀叹研究的质量低下,而不是用计算来掩盖它。第二项指控——藐视严格的条件——与第一项指控背道而驰;它可以建设性地解释为呼吁采用更新颖、更灵活和更有力的统计程序。该领域的领导者对基本科学方法和常识的健康关注,可能使他们在处理设计效果的量化和研究之间的非独立性建模等困难的统计问题时行动迟缓。但是,针对复杂任务量身定制的方法可能是可行的。所谓的“文件抽屉问题”的进展提供了一个先例。如果当研究发现影响在统计上不显著时,它们倾向于留在文件抽屉里,未发表,那么已发表研究的共识就倾向于夸大统计意义。罗森塔尔(Rosenthal)的“故障安全样本量”指数现在被广泛用于衡量这种危险,而艾扬格(Iyengar)和温室(Greenhouse)(4)则将最大似然估计的统计方法投入使用,以获得更精确的允许。一个公认的问题正在得到控制。元分析并不能避免困扰传统研究综合的任何严谨性问题。将在不同条件下对同一现象进行的实验结果结合起来,需要一个正式或非正式的模型,而模型的科学充分性始终是一个关键问题。评估模型相对于某些话语领域的充分性不是“程序”的问题。但是,从长远来看,统计实践中对系统误差来源的日益重视可能会有所贡献。其他这类问题也在议程上。医学荟萃分析发现,缺乏随机对照病例与治疗效果的估计强度之间存在相关性,这使研究设计的质量建模问题高度可见。独立的问题更令人生畏。忠诚网络和共同的先验信念不仅使独立研究团队的研究变得不那么独立,而且我们都赞赏的科学的累积特性也是如此。这里有一些相关的随机过程等待建模。迟早会有一个聪明的建模者接受挑战,并且会找到更合理的方法。然而,并不是我所有的朋友都会对更合理的方法感到满意。当一组研究过于混乱时,他们认为没有理由进行系统比较。我自己对这一指控的感受,受到了我回顾meta分析中最危险的练习之一的经历的影响。问题是学校废除种族隔离对黑人成绩测试成绩的短期影响。已知的研究有157项。有一些类似于控制设计的研究有19项。美国国家教育研究所委托六位持有不同观点的学者对这19项研究分别进行元分析。研究结果尚未发表,但哈里斯·库珀(Harris Cooper)发表了关于作者在进行分析时以及报告读者之间信念变化过程的研究。有机会亲自阅读这六份分析报告,我看着自己的信念发生了变化,并比较了库珀系统的、基于问卷调查的发现。在这些分析中,控制和可比性的问题一如既往地混乱。这些问题是否削弱了进行元分析的意义?在这19项研究中,约有三分之一的研究显示出了轻微的负面影响,其余的大多数显示出了轻微的积极影响,剩下的少数显示出了适度的积极影响。从研究报告中可以确定的设计和环境的不可比较的特征,在分析后,似乎只能很微弱地解释结果的差异,如果有的话。每项研究都有缺点,但缺点各不相同,我发现了一个相对随机的印象
i !OT LONG AGO AN ARTICLE IN Mercury CALLED "THE lunacy of it all" (1) examined published claims of correlations between manifestations of mental illness and phases of the moon. The authors J. Rotton and I. Kelly found different studies centering their statistically significant relationships on different lunar phases. "To deal with this problem," they wrote, "we resorted to meta-analysis, which is a statistical procedure for combining results from different studies. In our meta-analysis, we found no evidence for commonly held beliefs about the effects of a full Moon" (1, pp. 75 and 95). More and more scientists from all fields are resorting to "metaanalysis" when they review a body of scientific literature. Glass (2) coined the name "meta-analysis" in 1976; it designates research synthesis that uses formal statistical procedures to retrieve, select, and combine results from previous separate studies. A boom in meta-analyses is under way, but the boom is not being universally welcomed. I have friends who would gladly class meta-analysis itself among the fonms of lunacy whose relationship to lunar phases Rotton and Kelly were examining. Meta-analysis has waxed rapidly. Should we expect it to wane as rapidly, like the inconstant moon? My friends' skepticism centers, I think, around four charges. First is the suspicion that "garbage in and garbage out" are here being camouflaged by fancy statistics. Second is the view that an elementary mistake is being made when a potpourri of study outcomes is treated with procedures invented for a statistical sample of experimental outcomes, without the benefit of the controlled conditions, homogeneous measurement scales, and statistical independence that make the procedures valid. Third is a fear that the study of previous studies is being reduced to a routinized task ofcoding relegated to a research assistant, upping output per author-month by suppressing any role for wisdom. Related to this fear is a fourth feeling, that meta-analysis accommodates itself to a world in which bad science drives out good by weight of numbers. The first of these charges-the camouflage charge-is surely misdirected. There is very little fancy statistics at all in meta-analysis. The textbooks of the subject are exceedingly basic, down-to-earth, and opposed to all mystification. Examples, my favorites, are those by Hedges and Olkin, by Light and Pillemer, by Rosenthal, and by Wolf (3). Nor is there much tendency toward camouflage. Bad studies are exposed more baldly when one tries to extract or compute the size of an effect on a comparable basis study by study than when one relies, as informal research reviewers often do, on authors' own summaries of their conclusions. Authors of metaanalyses are forever deploring poor quality in studies rather than papering it over with calculations. The second charge-flouting the conditions for rigor-runs counter to the first; it could be interpreted, constructively, as a call for fancier, more flexible and robust statistical procedures. The healthy preoccupation of the leaders of the field with basic scientific method and common sense may have left them slow to take up hard statistical problems like the quantification of design effects and the modeling of nonindependence among studies. But methods tailored to the messiness of the task may well be possible. A precedent is provided by the progress on the so-called "file-drawer problem." If studies tend to remain in the file drawers, unpublished, when they find effects to be statistically insignificant, then the consensus of published studies will tend to exaggerate statistical significance. An index of Rosenthal's, the "fail-safe sample size," is now widely used to gauge this danger, while Iyengar and Greenhouse (4) have pressed statistical methods of maximum-likelihood estimation into service for more refined allowances. An admitted problem is being brought under control. Meta-analysis is exempt from none of the problems of rigor that beset traditional research synthesis. Combining results from experiments on the same phenomenon conducted under different conditions entails a model, formal or informal, and the scientific adequacy ofthe model is always a crucial question. Assessing the adequacy ofa model relative to some universe of discourse is not a matter for "procedures." But the increasing emphasis in statistical practice on sources of systematic error may have ideas to contribute in the long run. Other problems of this kind are on the agenda. Medical metaanalyses have turned up correlations between the absence of randomized controlled cases and the estimated strength of treatment effects, giving high visibility to the problem of modeling the quality of study designs. The issue of independence is more daunting. Not only do networks of loyalties and shared prior beliefs make studies by separate research teams less than independent, but so do the very cumulative properties of science that we all applaud. Some dependent stochastic process is here waiting to be modeled. Sooner or later an ingenious modeler is bound to take up the challenge, and sounder methods are going to be found. Not all my friends who press the second charge, however, are going to be satisfied by sounder methods. When a collection of studies is too messy, they see no reason for systematic comparisons. My own feelings about this charge have been affected by an experience of mine reviewing one of the diciest of all exercises in meta-analysis. The question was the short-term effect of school desegregation on blacks' achievement test scores. The known studies numbered 157. The studies with some semblance of controlled design numbered 19. The National Institute of Education commissioned six scholars with different prior views to do separate metaanalyses on these 19 studies. The results are unpublished, but Harris Cooper (5) has published studies of the process of belief change among the writers as they carried out their analyses and among readers of their reports. Given the chance to read the six analyses myself, I watched my own beliefs change and compared Cooper's systematic, questionnaire-based findings. The problems of control and comparability are as messy in these analyses as they can ever be. Do the problems vitiate the point of doing meta-analysis? About a third of the 19 studies had announced small negative effects, most of the rest small positive effects, and the remaining handful moderate positive effects. The noncomparable characteristics of design and context that could be identified from the study reports appeared, upon analysis, to account only tenuously, if at all, for the differences in outcome. Each ofthe studies had faults, but the faults differed, and I found the impression of a relatively random