From Dipsy-Doodle to Streaming Motions: Changes in Representation in the Analysis of Visual Scientific Data
From Dipsy-Doodle to Streaming Motions: Changes in Representation in the Analysis of Visual Scientific Data
复制标题
从 Dipsy-Doodle 到流媒体运动:视觉科学数据分析中表征的变化
DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
J. Trafton
中科院分区:
文献类型:
--
作者:
S. Trickett;W. Fu;Chrisitan D. Schunn;J. Trafton
From Dipsy-Doodles to Streaming Motions: Changes in Representation in the Analysis of Visual Scientific Data Susan B. Trickett Wai-Tat Fu Christian D. Schunn J. Gregory Trafton ( stricket@gmu.edu) ( wfu@gmu.edu) ( schunn@gmu.edu ) (trafton@itd.nrl.navy.mil) Department of Psychology George Mason University Fairfax, VA 22030 Department of Psychology George Mason University Fairfax, VA 22030 Naval Research Laboratory NRL Code 5513 Washington, DC 20375 Department of Psychology George Mason University Fairfax, VA 22030 Abstract This paper investigates the change in scientists' represen- tation of phenomena of interest during the exploratory analysis of visual data. The scientists initially represented expected findings in formal, scientific terms, whereas they represented anomalies in informal terms. Over time, these representations shifted from informal to formal. We pro- pose that this shift in representation is the result of an in- creased understanding of the individual phenomena, rather than of greater understanding of the data at a global level. Introduction A strong and perhaps foundational theme in cognitive sci- ence is the issue of representation. From both empirical and computational perspectives, performance has been found to depend heavily on how information is internally represented (Kotovsky, Hayes, & Simon, 1985; Larkin & Simon, 1987; Newell & Simon, 1972; Zhang & Norman, 1994). One area in which representation is likely to be especially important is scientific discovery (Schunn & Klahr, 1995). There are many formal and informal methods for represent- ing data, even within the same discipline and narrow sub- area. The choice of representation of the data is likely to have a large impact on what can and will be discovered. An additional twist on the issue of data representations in science is the difference between goals of scientific discovery and goals of communication of the discoveries. The external representations that are best for discovery are not necessarily the representations that are best for communication of the discovery to others. For example, issues of historical con- vention are likely to be more important in communication, whereas issues of ease of generation and manipulation are going to be more important for the original discovery. The goal of this paper is to examine how scientists repre- sent data internally to themselves while they are analyzing their data. In particular, do they tend to think of their data in formal, discipline-specific terms, or do they rely on more informal and simple perceptual terms? One might expect them to use formal terms because of their expertise and ex- tensive domain knowledge. On the other hand, they may use perceptual terms because in many areas of science, the data are presented in fairly complex visual displays that make heavy use of spatial metaphors—or indeed represent spatial dimensions directly (Trafton et al, under review). One dimension that we hypothesize would influence the choice of internal representation is the degree to which the data are as the scientist expects. That is, perhaps scientists are more likely to represent apparently anomalous data in informal, perceptual terms and expected data in formal, con- ceptual terms. Another related dimension that we investigated was time: How do scientists' representations of their data change over time as they explore their data? One might imagine that the representations become more formal as scientists develop an understanding of the dataset as a whole. Alternatively, the changes in representation may occur at a more item-specific level—the representation of each item changes separately as understanding of the item changes. A wide variety of methodologies has been used to study scientific reasoning and scientific discovery, each with their advantages and disadvantages (see Klahr & Simon, 1999, for a review). For this research project, we adopted a modified form of Kevin Dunbar's in vivo methodology (Dunbar, 1995, 1997, in press). The in vivo methodology involves observing scientists as they are doing their research. Dunbar focused on the activities that occur in lab group meetings. Because we were interested in the processes of data analysis, we focused, instead, on pairs of scientists working at their computers, analyzing their data. Like Dunbar, we perform a form of protocol analysis (Ericsson & Simon, 1993), ana- lyzing the speech produced by the scientists to make infer- ences about the underlying cognitive processes. The reason for focusing on pairs of scientists rather than on an individual scientist is that dyads produce speech natu- rally as part of their data analysis activities. By contrast, forcing an individual scientist to give a think-aloud protocol may change the very representations that we seek to study. For example, the individual scientist may change her focus to aspects of the data that are more easily verbalized, or she may change her representations from visio-spatial representa- tions to more verbal representations. Our methodology also contrasts with the retrospective analyses of historical cases from science (e.g., Gentner et al., 1997; Nersessian, 1985; Thagard, 1999). By focusing on the activities of non-famous (albeit expert) scientists work- ing on a problem that may or may not lead to an important discovery, we may obtain a more representative view of how scientists reason. 1 Of course, if one's goal is to understand how large conceptual leaps are made in science, the historical case-study approach may be more fruitful.