Disturbed by meta-analysis?
Disturbed by meta-analysis?
复制标题
受到荟萃分析的困扰?
作者:
K. Wachter
i !OT LONG AGO AN ARTICLE IN Mercury CALLED "THE lunacy of it all" (1) examined published claims of correlations between manifestations of mental illness and phases of the moon. The authors J. Rotton and I. Kelly found different studies centering their statistically significant relationships on different lunar phases. "To deal with this problem," they wrote, "we resorted to meta-analysis, which is a statistical procedure for combining results from different studies. In our meta-analysis, we found no evidence for commonly held beliefs about the effects of a full Moon" (1, pp. 75 and 95). More and more scientists from all fields are resorting to "metaanalysis" when they review a body of scientific literature. Glass (2) coined the name "meta-analysis" in 1976; it designates research synthesis that uses formal statistical procedures to retrieve, select, and combine results from previous separate studies. A boom in meta-analyses is under way, but the boom is not being universally welcomed. I have friends who would gladly class meta-analysis itself among the fonms of lunacy whose relationship to lunar phases Rotton and Kelly were examining. Meta-analysis has waxed rapidly. Should we expect it to wane as rapidly, like the inconstant moon? My friends' skepticism centers, I think, around four charges. First is the suspicion that "garbage in and garbage out" are here being camouflaged by fancy statistics. Second is the view that an elementary mistake is being made when a potpourri of study outcomes is treated with procedures invented for a statistical sample of experimental outcomes, without the benefit of the controlled conditions, homogeneous measurement scales, and statistical independence that make the procedures valid. Third is a fear that the study of previous studies is being reduced to a routinized task ofcoding relegated to a research assistant, upping output per author-month by suppressing any role for wisdom. Related to this fear is a fourth feeling, that meta-analysis accommodates itself to a world in which bad science drives out good by weight of numbers. The first of these charges-the camouflage charge-is surely misdirected. There is very little fancy statistics at all in meta-analysis. The textbooks of the subject are exceedingly basic, down-to-earth, and opposed to all mystification. Examples, my favorites, are those by Hedges and Olkin, by Light and Pillemer, by Rosenthal, and by Wolf (3). Nor is there much tendency toward camouflage. Bad studies are exposed more baldly when one tries to extract or compute the size of an effect on a comparable basis study by study than when one relies, as informal research reviewers often do, on authors' own summaries of their conclusions. Authors of metaanalyses are forever deploring poor quality in studies rather than papering it over with calculations. The second charge-flouting the conditions for rigor-runs counter to the first; it could be interpreted, constructively, as a call for fancier, more flexible and robust statistical procedures. The healthy preoccupation of the leaders of the field with basic scientific method and common sense may have left them slow to take up hard statistical problems like the quantification of design effects and the modeling of nonindependence among studies. But methods tailored to the messiness of the task may well be possible. A precedent is provided by the progress on the so-called "file-drawer problem." If studies tend to remain in the file drawers, unpublished, when they find effects to be statistically insignificant, then the consensus of published studies will tend to exaggerate statistical significance. An index of Rosenthal's, the "fail-safe sample size," is now widely used to gauge this danger, while Iyengar and Greenhouse (4) have pressed statistical methods of maximum-likelihood estimation into service for more refined allowances. An admitted problem is being brought under control. Meta-analysis is exempt from none of the problems of rigor that beset traditional research synthesis. Combining results from experiments on the same phenomenon conducted under different conditions entails a model, formal or informal, and the scientific adequacy ofthe model is always a crucial question. Assessing the adequacy ofa model relative to some universe of discourse is not a matter for "procedures." But the increasing emphasis in statistical practice on sources of systematic error may have ideas to contribute in the long run. Other problems of this kind are on the agenda. Medical metaanalyses have turned up correlations between the absence of randomized controlled cases and the estimated strength of treatment effects, giving high visibility to the problem of modeling the quality of study designs. The issue of independence is more daunting. Not only do networks of loyalties and shared prior beliefs make studies by separate research teams less than independent, but so do the very cumulative properties of science that we all applaud. Some dependent stochastic process is here waiting to be modeled. Sooner or later an ingenious modeler is bound to take up the challenge, and sounder methods are going to be found. Not all my friends who press the second charge, however, are going to be satisfied by sounder methods. When a collection of studies is too messy, they see no reason for systematic comparisons. My own feelings about this charge have been affected by an experience of mine reviewing one of the diciest of all exercises in meta-analysis. The question was the short-term effect of school desegregation on blacks' achievement test scores. The known studies numbered 157. The studies with some semblance of controlled design numbered 19. The National Institute of Education commissioned six scholars with different prior views to do separate metaanalyses on these 19 studies. The results are unpublished, but Harris Cooper (5) has published studies of the process of belief change among the writers as they carried out their analyses and among readers of their reports. Given the chance to read the six analyses myself, I watched my own beliefs change and compared Cooper's systematic, questionnaire-based findings. The problems of control and comparability are as messy in these analyses as they can ever be. Do the problems vitiate the point of doing meta-analysis? About a third of the 19 studies had announced small negative effects, most of the rest small positive effects, and the remaining handful moderate positive effects. The noncomparable characteristics of design and context that could be identified from the study reports appeared, upon analysis, to account only tenuously, if at all, for the differences in outcome. Each ofthe studies had faults, but the faults differed, and I found the impression of a relatively random