Individual Submission Summary
Share...

Direct link:

What Changes Across the Measurement of Classroom Quality During High- and Low-Stakes Assessments?

Wed, April 7, 4:30 to 5:30pm EDT (4:30 to 5:30pm EDT), Virtual

Abstract

Quality Rating and Improvement Systems (QRISs) across the U.S. are varied and distinct, yet one common element is often the measurement of quality within childcare centers and family childcare homes. The length of time between assessments and the number of classrooms assessed within a program varies widely across states (QRIS Compendium, 2019). One of the key challenges of quality ratings is the assumption that the ratings on the day of the ‘official’ assessment is representative of quality on other days. However, inconsistency in quality may have important implications for children’s experiences and may shed light on challenges of associating classroom quality with children’s developmental outcomes in the short and long-term. Bioecological theory of human development (Bronfenbrenner & Morris, 1988; 2006) highlights the value of consistent and progressively more complex reciprocal interactions with people, symbols, and objects within daily environments on children’s development. If early childhood environments are inconsistent in quality across official (high stakes) and unofficial (low stakes) assessments this may have implications for the set up of QRISs.
This study investigated whether quality changes between low and high-stakes early childhood classroom assessments. A paired sample of 23 Early Childhood Environment Rating Scale-Revised (ECERS-R; Harms et al., 2005) assessments and a matched sample of 153 Infant Toddler Environment Rating Scale-Revised (ITERS-R; Harms et al., 2006) assessments were analyzed. The number of children present and the number of months the teacher had been in the classroom assessed were included as co-variates in the analyses. All assessments were completed by highly trained state-level assessors.
Comparative analyses were run across the high- and low-stakes assessments. The ECERS-R data were paired comparisons and the ITERS-R were independent sample comparisons. Analyses examined the overall average scores, the subscales scores, previously established factor-based scores for the ECERS-R (Cassidy et al., 2005) and ITERS-R (Hestenes et al., 2007), and item-level differences. The results for the ECERS-R showed significantly lower scores on the low-stakes compared to the high-stakes assessments for the average score, the Personal Care Routines and Interaction subscales, and the Language/Interaction factor-based score (Table 1). Items with significantly lower scores occurred for nap/rest, health practices, supervision of gross motor activities, discipline, and staff-child interactions. The ITERS-R results showed significantly lower scores for the low-stakes assessments compared to the high-stakes for the average score, all six subscales, the three factor-based scores, and almost all the items (Table 2). Although the paired samples for the ECERS-R were relatively small, and the ITERS-R comparisons were based on a matched sample, the results show a consistent pattern of quality dropping after an official assessment. These findings strengthen a growing concern in the field that classrooms may not maintain continuous quality over time. Children’s daily experiences appear to not match the level of quality determined by high stakes assessments. States may need to consider more frequent checks on quality, unannounced quality assessments, or more strategic efforts within technical assistance support to help programs maintain levels of quality over time.

Authors