From Dipsy-Doodle to Streaming Motions: Changes in Representation in the Analysis of Visual Scientific Data

From Dipsy-Doodle to Streaming Motions: Changes in Representation in the Analysis of Visual Scientific Data
复制标题

从 Dipsy-Doodle 到流媒体运动:视觉科学数据分析中表征的变化

DOI:
--
复制
发表时间:
2000
期刊:
--
影响因子:
--
通讯作者:
J. Trafton
J. Trafton
中科院分区:
--
文献类型:
--
作者:
S. Trickett;W. Fu;Chrisitan D. Schunn;J. Trafton

文献摘要

被引文献

相似文献

从Dipsy-Doodles到Stream Motion:视觉科学数据分析中表征的变化Susan B.Trickett Wai-tat Fu Christian D.Schunn J.Gregory Trafton(Strecket@gmu.edu)(wfu@gmu.edu)(schunn@gmu.edu)(trafton@itd.nrl.navy.mil)乔治梅森大学费尔法克斯,弗吉尼亚州22030心理学系乔治梅森大学费尔法克斯,VA 22030海军研究实验室乔治梅森大学费尔法克斯VA 22030摘要本文考察了在视觉数据的探索性分析过程中,科学家对感兴趣现象的表征的变化。科学家最初用正式的、科学的术语表示预期的发现,而他们用非正式的术语表示异常。随着时间的推移,这些陈述从非正式转变为正式。我们认为,这种代表性的转变是对个别现象的理解增加的结果,而不是在全球层面对数据的更好理解的结果。引言认知科学中一个强有力的、或许是基础性的主题是表征问题。从经验和计算的角度来看,绩效在很大程度上取决于信息的内部表征(Kotovsky,Hayes,&Simon,1985;Larkin&Simon,1987;Newell&Simon,1972;Zhang&Norman,1994)。代表性可能特别重要的一个领域是科学发现(Schunn&Klahr,1995)。有许多正式和非正式的方法来表示数据,即使在相同的学科和狭窄的子领域内也是如此。数据表示形式的选择可能会对能够和将发现什么产生很大影响。关于科学中的数据表示问题的另一个转折是科学发现的目标和发现的传播目标之间的差异。最适合于发现的外部表示不一定是最适合将发现传达给他人的表示。例如,历史发明的问题在交流中可能更加重要,而对于原始发现来说,易用性和操纵性问题将更加重要。这篇论文的目的是研究科学家在分析他们的数据时如何在内部向他们自己发送数据。特别是,他们是倾向于以正式的、特定于学科的术语来考虑他们的数据,还是依赖于更非正式和简单的感性术语?人们可能会认为他们会使用正式的术语,因为他们的专业知识和广泛的领域知识。另一方面,他们可能使用知觉术语,因为在许多科学领域,数据以相当复杂的视觉显示方式呈现,大量使用空间隐喻--或者实际上直接表示空间维度(Trafton等人,正在审查中)。我们假设会影响内部表示法选择的一个维度是数据达到科学家预期的程度。也就是说,也许科学家更有可能用非正式的、感性的术语来表示明显异常的数据,而用正式的、概念的术语来表示预期的数据。我们调查的另一个相关维度是时间:当科学家探索他们的数据时,他们对数据的表示如何随着时间的推移而变化?人们可能会认为,随着科学家对数据集作为一个整体的理解,这些表示法会变得更加正式。或者,表示的改变可以发生在更特定于项的级别上--随着对项的理解的改变,每个项的表示单独地改变。各种各样的方法被用来研究科学推理和科学发现,每一种方法都有它们的优点和缺点(参见Klahr&Simon,1999)。对于这个研究项目,我们采用了Kevin Dunbar体内方法学的修改形式(Dunbar,1995,1997,出版)。体内方法论包括在科学家进行研究的同时观察他们。邓巴把重点放在了实验室小组会议上的活动上。因为我们对数据分析过程感兴趣,所以我们关注的是两个科学家在他们的电脑前工作,分析他们的数据。就像邓巴一样,我们进行一种形式的协议分析(Ericsson&Simon,1993),对科学家产生的言语进行分析,以推断潜在的认知过程。关注成对科学家而不是单个科学家的原因是,作为数据分析活动的一部分,双星会产生言语自然集会。相比之下,强迫一位科学家给出一个有声思考的方案可能会改变我们试图研究的陈述。例如,单个科学家可能会将她的重点转向更容易用语言表达的数据方面,或者她可能会将她的表示从视觉空间表示改变为更多的语言表示。我们的方法也与对科学的历史案例的回顾分析形成对比(例如,Gentner等人,1997;Nersesian,1985;Thagard,1999)。通过关注非著名(尽管是专家)科学家在一个可能导致或可能不会导致重要发现的问题上的活动,我们可能会对科学家如何推理获得更具代表性的观点。当然,如果一个人的目标是了解科学在概念上取得了多大的飞跃,历史案例研究方法可能会更有成效。
From Dipsy-Doodles to Streaming Motions: Changes in Representation in the Analysis of Visual Scientific Data Susan B. Trickett Wai-Tat Fu Christian D. Schunn J. Gregory Trafton ( stricket@gmu.edu) ( wfu@gmu.edu) ( schunn@gmu.edu ) (trafton@itd.nrl.navy.mil) Department of Psychology George Mason University Fairfax, VA 22030 Department of Psychology George Mason University Fairfax, VA 22030 Naval Research Laboratory NRL Code 5513 Washington, DC 20375 Department of Psychology George Mason University Fairfax, VA 22030 Abstract This paper investigates the change in scientists' represen- tation of phenomena of interest during the exploratory analysis of visual data. The scientists initially represented expected findings in formal, scientific terms, whereas they represented anomalies in informal terms. Over time, these representations shifted from informal to formal. We pro- pose that this shift in representation is the result of an in- creased understanding of the individual phenomena, rather than of greater understanding of the data at a global level. Introduction A strong and perhaps foundational theme in cognitive sci- ence is the issue of representation. From both empirical and computational perspectives, performance has been found to depend heavily on how information is internally represented (Kotovsky, Hayes, & Simon, 1985; Larkin & Simon, 1987; Newell & Simon, 1972; Zhang & Norman, 1994). One area in which representation is likely to be especially important is scientific discovery (Schunn & Klahr, 1995). There are many formal and informal methods for represent- ing data, even within the same discipline and narrow sub- area. The choice of representation of the data is likely to have a large impact on what can and will be discovered. An additional twist on the issue of data representations in science is the difference between goals of scientific discovery and goals of communication of the discoveries. The external representations that are best for discovery are not necessarily the representations that are best for communication of the discovery to others. For example, issues of historical con- vention are likely to be more important in communication, whereas issues of ease of generation and manipulation are going to be more important for the original discovery. The goal of this paper is to examine how scientists repre- sent data internally to themselves while they are analyzing their data. In particular, do they tend to think of their data in formal, discipline-specific terms, or do they rely on more informal and simple perceptual terms? One might expect them to use formal terms because of their expertise and ex- tensive domain knowledge. On the other hand, they may use perceptual terms because in many areas of science, the data are presented in fairly complex visual displays that make heavy use of spatial metaphors—or indeed represent spatial dimensions directly (Trafton et al, under review). One dimension that we hypothesize would influence the choice of internal representation is the degree to which the data are as the scientist expects. That is, perhaps scientists are more likely to represent apparently anomalous data in informal, perceptual terms and expected data in formal, con- ceptual terms. Another related dimension that we investigated was time: How do scientists' representations of their data change over time as they explore their data? One might imagine that the representations become more formal as scientists develop an understanding of the dataset as a whole. Alternatively, the changes in representation may occur at a more item-specific level—the representation of each item changes separately as understanding of the item changes. A wide variety of methodologies has been used to study scientific reasoning and scientific discovery, each with their advantages and disadvantages (see Klahr & Simon, 1999, for a review). For this research project, we adopted a modified form of Kevin Dunbar's in vivo methodology (Dunbar, 1995, 1997, in press). The in vivo methodology involves observing scientists as they are doing their research. Dunbar focused on the activities that occur in lab group meetings. Because we were interested in the processes of data analysis, we focused, instead, on pairs of scientists working at their computers, analyzing their data. Like Dunbar, we perform a form of protocol analysis (Ericsson & Simon, 1993), ana- lyzing the speech produced by the scientists to make infer- ences about the underlying cognitive processes. The reason for focusing on pairs of scientists rather than on an individual scientist is that dyads produce speech natu- rally as part of their data analysis activities. By contrast, forcing an individual scientist to give a think-aloud protocol may change the very representations that we seek to study. For example, the individual scientist may change her focus to aspects of the data that are more easily verbalized, or she may change her representations from visio-spatial representa- tions to more verbal representations. Our methodology also contrasts with the retrospective analyses of historical cases from science (e.g., Gentner et al., 1997; Nersessian, 1985; Thagard, 1999). By focusing on the activities of non-famous (albeit expert) scientists work- ing on a problem that may or may not lead to an important discovery, we may obtain a more representative view of how scientists reason. 1 Of course, if one's goal is to understand how large conceptual leaps are made in science, the historical case-study approach may be more fruitful.