Individual Submission Summary
Share...

Direct link:

Reliability and Validity of the Classroom Assessment Scoring System, Pre-K Using Full-Day/Full-Week Video Streams

Thu, April 8, 1:10 to 2:40pm EDT (1:10 to 2:40pm EDT), Virtual

Abstract

The preschool version of the Classroom Assessment Scoring System (CLASS PreK) is widely used in research, practice and policy. Yet, concerns are increasingly raised about its low inter-rater reliability, ample occasion variance, and wide error bands, especially when stakes are high (Burchinal, 2018; Mashburn, 2017). Although empirical evidence about these concerns is growing, relatively few studies have looked at these issues for the PreK version of CLASS.
Our manuscript offers new evidence using video of 425 15-minute occasions taken from 58 days within 10 preschool-aged classrooms serving children from diverse racial-ethnic backgrounds. Fourteen raters were trained by a certified CLASS trainer and passed the Teachstone CLASS PreK reliability test for the 10 items from 3 domains. These raters scored occasions in a partially crossed design that achieved connectedness by having all raters score some occasions and pairs or triplets of raters score other occasions.
We used these data to examine: (1) What is the level of inter-rater reliability? (2) How much variation is within and between classrooms? (3) How precise are estimates? Going beyond prior studies, we demonstrate for substantive scholars and non-academic audiences the utility of the many-facet Rasch model (MFRM) to complement traditional statistics (illustrative graphs in Figures 1 and 2).
First, our estimates quantify how Teachstone’s within-one certification criterion allows considerable rater variation. Although within-one agreement was high for our certified raters (69% to 96%; median = 77%), exact agreement was lower (26% to 64%; median = 33%), importantly replicating the limited published exact agreement statistics for CLASS PreK (Cash, Hamre, Pianta, & Myers, 2012). Kappa statistics were likewise moderate or lower (median = .07 exact, .32 within-one). A g-study further attributed from one-fifth to one-third of the variance to a rater main effect, with another third to half reflecting interactions across facets and residual variance. Extending prior work, the MFRM reinforced how over half the raters differed significantly in their average harshness, being reliably separable into 3-to-7 harshness levels across the three domains. To correct for such systematic rater effects (i.e., scores depend on who conducted the observation), we demonstrate how the MFRM can produce adjusted scores and discuss the circumstances under which research, practice, and policy could use such adjustments.
Second, occasion variance was also considerable. The g-study apportioned about one-third of variance to an occasion main effect across domains, highlighting the ways in which a classroom’s score depends on when an observation was conducted. Third, the MFRM also identified that the items were not ideally targeted at sampled occasions. Even though estimates spanned 8 to 9 points in the MFRM logit metric, the true score variance was just twice the error variance, and, occasions could be reliably separated into just 3 levels (e.g., low, moderate, and high quality). We demonstrate how the MFRM can be used to foreground the ways occasions can be considered conceptually nested within classrooms as well as to achieve continuous measure improvement by visually arraying items alongside occasions and by analyzing how scores fluctuate across daily routines and activities.

Authors