Improving STEM Program Quality in Out-of-School-Time: Tool Development and Validation.
Improving STEM Program Quality in Out-of-School-Time: Tool Development and Validation.
复制标题
提高课外时间的 STEM 项目质量:工具开发和验证。
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
G. Noam
中科院分区:
文献类型:
--
作者:
A. Shah;C. Wylie;D. Gitomer;G. Noam
In and out-of-school time (OST) experiences are viewed as complementary in contributing to students’ interest, engagement, and performance in science, technology, engineering, and mathematics (STEM). While tools exist to measure quality in general afterschool settings and others to measure structured science classroom experiences, there is a need for reliable measures of STEM program qual-ityinOSTsettingssuchasafterschoolprograms,summercamps,and museum or science center programming. In this paper we present the development of the Dimensions of Success (DoS) tool, which defines twelve key components of informal, exploratory STEM programming that goes beyond the school day. Additionally, we present avalidityargumentthatincludesreliabilityevidencefortheDoStool based on two studies: Study 1 ( n = 284 observations) and Study 2 ( n = 56 observations). Our findings suggest that the coherence of the constructs and validity evidence, as well as the training and certification procedures in place for DoS, make it an important tool to understand the quality of STEM experiences for youth beyond the school day. A tool like DoS has several implications, including the ability to make national comparisons across programs, create aggre-gate databases, improve program quality and professional development, as well as to link program quality to student-level outcomes. G-coefficient was used to estimate dimension reliabilities after two observations by two observers (see Appendix). We were then able to conduct a dependability (D) study to estimate how many observations would be needed to get a more reliable estimate of the quality of a module (keeping two observers constant). These analyses (Appendix) indicated that multiple observations are needed to get a stable measure of quality and that the most stable measure was to use the two-factor structure that was identified from the Exploratory Factor Analysis to create two composite scores for the quality of the learning environment and the quality of STEM meaning-making. Interestingly, more observations were needed to understand the STEM meaning-making factor versus the learning environment factors: While four observations would likely result in a reliable estimate of the learning environment factor, even 10 observations would be insufficient for the STEM content factor, given the levels of inter-observer agreement in the current study. Given these findings, we made changes to the training, certification, and calibration process and followed up with Study 2 to examine the impact on resulting assessor reliability.