Shifting sights on STEM education quantitative instrumentation development: The importance of moving validity evidence to the forefront rather than a footnote

Shifting sights on STEM education quantitative instrumentation development: The importance of moving validity evidence to the forefront rather than a footnote
复制标题

转变对 STEM 教育定量仪器开发的关注:将有效性证据置于最前沿而不是脚注的重要性

DOI:
10.1111/ssm.12410
复制
发表时间:
2020
影响因子:
1.1
通讯作者:
Sondergeld, Toni A.
Sondergeld, Toni A.
中科院分区:
--
文献类型:
--
作者:
Sondergeld, Toni A.

文献摘要

参考文献

被引文献

相似文献

一个人很难完成教育学的研究生课程,至少不会听到有效性(衡量我们想要衡量的东西)和可靠性(始终如一地衡量)的概念。作为教育研究人员,我们很清楚,我们的工具的强有力的有效性和可靠性证据是研究进步的基础。即便如此,硬科学家通常认为教育研究较少,因为“社会科学是,嗯,'软',缺乏方法论的严谨性”(“赞美软科学”,2005年,第2页)。这种批评在一定程度上是由于所研究的结构的性质。虽然我们的硬科学同事正在使用成熟的工具来研究科学现象,以测量身高增长或速度差异等结果;社会科学研究人员正试图量化或解释更具挑战性的测量结构,如人类的态度和信仰,或认知能力。这一挑战也可能部分是由于缺乏测量培训(Liu,2010; Shih,Reys,Reys,& Engledowl,2019; Smith,Conrad,Chang,& Piazza,2002),这对社会科学家开发和验证声音工具至关重要。为了解决社会科学中对仪器严谨性的固有担忧,教育和心理测试标准于1966年首次由美国教育研究协会(AERA),美国心理协会(阿帕)和国家教育测量理事会(NCME)合作发布。2014年,最新版本的标准(AERA,阿帕和NCME,2014)讨论了需要收集,评估和记录多种形式的有效性证据,以确定定量工具的结果和解释是否适合特定的意图。此外,随着更大的有效性证据来支持仪器的有效性论证,可以得出关于仪器可靠性的更强的推论(AERA等人,2014; Kane,2016)。虽然在大量的文献中讨论了许多形式的有效性证据,但有五种具体类型。标准敦促教育和心理评估的开发者评估:测试内容,反应过程,内部结构,与其他变量的关系,以及后果(AERA等人,2014年)。测试内容效度证据调查工具项目对齐(测试内容)与被测结构(理论特质)。支持证据通常来自主题专家(SME)评估项目到结构的对齐,可以是逻辑的或经验的(定性的)(Sireci & Faulkner-Bond,2014)。响应过程有效性证据评估参与者的响应或与测试结构的性能一致性(Leighton,2017)。一般来说,支持响应过程有效性的数据是定性的,并通过认知访谈收集,大声思考,或与典型受访者的样本进行焦点小组访谈,以检查他们是否理解项目并以开发人员设想的方式进行响应(帕迪利亚&贝尼特斯,2014)。通过心理测量学方法评估内部结构效度证据,以探索:(a)工具维度,(B)测量不变性,(c)工具可靠性(Rios &威尔斯,2014)。与其他变量的关系效度证据通常使用统计检验来调查工具结果与其他假设相关的变量的关联(无论是积极的还是消极的)(Beckman,Cook,& Mandrekar,2005)。结果有效性证据和偏见通常通过定性数据来收集,以检查参与者如何看待...
One would be hard pressed to complete a graduate program in education and not, at a minimum, hear about the concepts of validity (measuring what we intend to measure) and reliability (measuring consistently). As educational researchers, we are well aware that strong validity and reliability evidence for our instruments are fundamental for the advancement of research. Even so, hard scientists often think less of educational research because “the social sciences are, well,‘soft’, and lacking methodological rigour”(“In praise of soft science,” 2005, p. 2). This criticism is in part due to the nature of the constructs under study. While our hard science colleagues are investigating scientific phenomena with well-established tools to measure outcomes such as growth in height or differences in speed; social science researchers are attempting to quantify or explain constructs that are far more challenging to measure such as human attitudes and beliefs, or cognitive abilities. The challenge may also in part be due to a lack of measurement training (Liu, 2010; Shih, Reys, Reys, & Engledowl, 2019; Smith, Conrad, Chang, & Piazza, 2002), which is essential for social scientists to develop and validate sound instruments. To address this inherent concern about instrumentation rigor in the social sciences, The Standards for Educational and Psychological Testing were first released in 1966 by a collaboration comprised of American Educational Research Association (AERA), American Psychological Association (APA), and National Council on Measurement in Education (NCME). In 2014, the most recent version of The Standards (AERA, APA, & NCME, 2014) discussed the need to collect, evaluate, and document multiple forms of validity evidence for the results and interpretations of a quantitative instrument to be judged suitable for a specified intent. Further, with greater validity evidence to support the validity argument of an instrument, stronger inferences may be drawn regarding instrumentation soundness (AERA et al., 2014; Kane, 2016). While there are numerous forms of validity evidence discussed within volumes of literature, there are five specific types The Standards urge developers of educational and psychological assessments to evaluate: test content, response processes, internal structure, relationship to other variables, and consequential (AERA et al., 2014). Test content validity evidence investigates instrument item alignment (test content) with the construct to be measured (theoretical trait). Supporting evidence often comes from subject matter experts (SMEs) evaluating item-to-construct alignment and can be logical or empirical (qualitative)(Sireci & Faulkner-Bond, 2014). Response process validity evidence appraises participant responses or performance alignment with the test construct (Leighton, 2017). Generally, data to support response process validity is qualitative and collected through cognitive interviews, think alouds, or focus group interviews with a sample of typical respondents to check that they understand items and respond in ways developers envisioned (Padilla & Benitez, 2014). Internal structure validity evidence is assessed through psychometric methods to explore:(a) instrument dimensionality,(b) measurement invariance, and (c) instrument reliability (Rios & Wells, 2014). Relationship to other variables validity evidence often uses statistical testing to investigate instrument outcome associations with other variables hypothesized to be related (either positively or negatively)(Beckman, Cook, & Mandrekar, 2005). Consequential validity evidence and bias are often collected through qualitative data to examine how participants perceive the …
在教育研究中使用有声思考访谈和认知实验室
DOI: --
发表时间: 2017
期刊:
影响因子: --
作者:
Jacqueline P. Leighton
通讯作者: Jacqueline P. Leighton
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
Jeffrey C. Shih;Robert E. Reys;Barbara J. Reys;Christopher Engledowl
通讯作者: Christopher Engledowl
基于 TRILHAR 的反应过程的有效性证据 - 婴儿词汇筛选工具。
DOI: --
发表时间: 2021
期刊: International Symposium on Cooperative Database Systems for Advanced Applications
影响因子: --
作者:
Alexandre Lucas de Araújo Barbosa;Cíntia Alves Salgado Azoni
通讯作者: Cíntia Alves Salgado Azoni
基于测试内容的有效性证据。
DOI: --
发表时间: 2014
期刊: Psicothema
影响因子: 3.6
作者:
S. Sireci;Molly Faulkner
通讯作者: Molly Faulkner
DOI: 10.1111/j.1525-1497.2005.0258.x
发表时间: 2005-12-01
影响因子: 5.7
作者:
Beckman, TJ;Cook, DA;Mandrekar, JN
通讯作者: Mandrekar, JN